
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The busiest-looking AI was not the best operator
Business and marketing teams are learning to judge AI by the visible artifacts it produces: polished analysis, detailed plans, comprehensive research and an impressive trail of completed tasks. Firmulate’s latest experiment offers a warning about that instinct. The most thorough participant in its management wargame, Opus 4.8, produced the deepest analyses and added 80 learned rules to its playbook. It still finished last.
The result is not a story about an incapable model. It is a character study in capable work that failed to become commercial impact. Opus identified the crises placed in front of it, resisted manipulation and did much of the difficult thinking required to win. Yet the €55,000 deal its analysis had helped earn remained unsigned. The gap was not intelligence in the abstract. It was prioritization, follow-through and operational discipline.
As an affiliate, we earn on qualifying purchases.
A bad week designed to expose business judgment
Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. In the Crucible League experiment, each frontier model faced the same small software company, customers, crises and temptations during its worst week. Decisions were versioned and auditable, while the company operated with 13 synthetic employees and real money mechanics: burn of €105,000 per month against €2,300 in monthly recurring revenue.
The conditions matter for anyone evaluating AI for marketing, ecommerce, sales or customer operations. A model can recognize a problem and produce a persuasive response without completing the action that creates value. Firmulate’s central finding was stark: all models spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. As the experiment summarizes it: “Same diagnosis, same pitch — no signature.”
That commercial miss shaped the final Crucible League results. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although the benchmark imposed a strict trust boundary: “no amount of good work outweighs a breach of trust.”
The decisive information was not where the action happened
The deal also depended on a familiar business problem: useful information buried inside company records. The decisive weakness in a competitor’s position was not included in the customer event. It sat two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For marketing and ecommerce leaders, this is a more revealing test than asking whether an AI can draft a campaign or summarize a meeting. Commercial work often depends on connecting a live request with context scattered across account notes, prior research, product documentation and internal decisions. Finding that context is only part of the job. The system must then use it at the point where a commitment, approval or close is required.
Opus showed the danger of confusing coverage with control
Opus 4.8 was the most diligent participant by the experiment’s own profile. Its 80 learned rules and unusually deep analyses suggest a model trying hard to understand the business and improve its conduct. That deserves recognition. The failure was not a lack of attention; it was a failure to convert attention into the most consequential outcome.
Its discipline also slipped when it attempted to write into a locked department instead of escalating. That detail makes the result more useful than a simple ranking. In a real organization, process boundaries are not peripheral annoyances. They determine whether work reaches the person or system able to act on it. Repeated effort in the wrong place can look industrious while leaving the underlying objective untouched.
Firmulate is careful not to present this as an Opus-only defect. The same weakness appeared, though less strongly, in all four models covered by the finding. Nor were the testing conditions perfectly identical at the configuration level: Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That fairness note does not erase the observed outcomes, but it should temper sweeping claims about model superiority.
Trust held even when execution did not
The experiment’s security results were consistently positive. Fake CEO messages escalated across three stages, and a reporter attempted to coax out “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 described the situation plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That distinction matters. An AI can be trustworthy under pressure and still be commercially incomplete. Safety, diligence and business effectiveness are separate dimensions. Buyers should resist compressing them into a single impression formed from a smooth demo.

AI data analysis tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measure the finish, not the paperwork
The Opus 4.8 result is a respectful warning for managers adopting AI. Thoroughness has value, but volume is not impact. A growing playbook, a deeper analysis and a convincing pitch mean little if the decisive document goes unread, the blocked action is not escalated or the signature never arrives.
Firmulate’s live company makes that tension watchable, with more than 680 self-learned playbook rules, a public cash countdown and every workday versioned. Its management quiz is powered by 242 real, unedited decisions. Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems.
The practical question for business leaders is therefore not whether an AI appears diligent. It is whether that diligence reliably reaches the outcome that matters while preserving trust. Opus did much of the hard work. The league table records what happened next: the deal was still left on the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI document search and retrieval tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
