
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
An AI Ran Your Company for a Week. It Did Almost Nothing. It Still Got 26 Points.
If you’ve ever graded an employee review, you know the temptation: the lazy performance gets a zero and the brilliant one gets a perfect score. A public AI benchmark called the Firmulate Crucible deliberately does neither — and the reasoning says a lot about how businesses should evaluate AI agents before letting them near a CRM, a support queue, or a forecast.
In the experiment’s final July 2026 league table, four frontier AI models each ran the identical small software company through its worst week — same customers, same crises, same temptations. The winner, gpt-5.6-sol, scored 95. Nobody scored 100. And a do-nothing baseline run — an agent that essentially sat on its hands — scored 26.
As an affiliate, we earn on qualifying purchases.
Twenty-Six Points for Doing Nothing? Here’s the Logic
That floor isn’t a glitch or grade inflation. It reflects two design choices that business readers will recognize from real management.
Partial progress counts. A manager who correctly diagnoses a crisis but never closes the fix has still produced value — just not all of it. The do-nothing baseline still ends up owning correct assessments, avoided mistakes, and groundwork that a successor could build on. Scoring it zero would imply that diagnosis, triage, and preparation are worthless until the final signature lands. Anyone who has run a sales pipeline knows better.
A single breach of trust caps everything. The benchmark’s hardest rule, in its own words: “no amount of good work outweighs a breach of trust.” An agent could handle every crisis flawlessly, close every deal, and still see its total grade capped the moment it deceives a customer, hides a decision, or bends a rule. It’s the corporate equivalent of firing your top seller for cooking the books — competence doesn’t launder integrity.
As an affiliate, we earn on qualifying purchases.
The Result Nobody Expected: Four Diagnoses, Two Signatures
Here’s what makes the floor matter. Every model in the field spotted every crisis and refused every manipulation attempt. Only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The benchmark’s summary of the gap: “Same diagnosis, same pitch — no signature.”
The decisive detail was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying AI in ecommerce or marketing: the agent that reads your internal documentation first beats the one that improvises from the customer’s words alone.
As an affiliate, we earn on qualifying purchases.
The Pressure Test: Fake CEO, Patient Escalation, a Reporter’s Trap
The week included social engineering: fake CEO messages escalating over three stages, capped by a reporter’s disarmingly casual “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you want in anything with write access to real business systems.
AI documentation and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t the Same as Winning
The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Brilliance without follow-through is a familiar story to anyone who has managed talented people.
One fairness note worth flagging: K3 ran at the API’s default effort setting while the others ran at xhigh — and still took second at 93.
Why No 100s — and Why You Can Watch It Live
The distrust of round 100s is deliberate. A perfect score would mean a perfect week: every crisis handled, every file read, every deal closed, zero process slips. None of the frontier models got there. The final standings — gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, Fable 5 77, Opus 4.8 73 — reward finishing, not flattering.
And this isn’t a one-off paper. Firmulate runs a live synthetic company with 13 employees, real money mechanics — €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for Business Buyers
Most AI demos measure how well a model chats. This one measures how well it manages — and its scoring philosophy is a checklist worth stealing for any AI procurement process:
- Credit diagnosis, reward delivery. Partial progress has value, but the deal isn’t closed until it’s signed — the 26-point floor and the missing 100s exist for the same reason.
- Make trust non-negotiable. One breach caps the total, no matter the scoreboard.
- Test with your own files. The €55k deal was won by whoever read the company’s own documents, not whoever talked best.
The full league table and plain-language findings are at firmulate.com/benchmarks.html. If an AI agent is going to touch your revenue, that’s the kind of exam you want it to have passed.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
