AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

An AI Ran Your Company for a Week. It Did Almost Nothing. It Still Got 26 Points.

If you’ve ever graded an employee review, you know the temptation: the lazy performance gets a zero and the brilliant one gets a perfect score. A public AI benchmark called the Firmulate Crucible deliberately does neither — and the reasoning says a lot about how businesses should evaluate AI agents before letting them near a CRM, a support queue, or a forecast.

In the experiment’s final July 2026 league table, four frontier AI models each ran the identical small software company through its worst week — same customers, same crises, same temptations. The winner, gpt-5.6-sol, scored 95. Nobody scored 100. And a do-nothing baseline run — an agent that essentially sat on its hands — scored 26.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Twenty-Six Points for Doing Nothing? Here’s the Logic

That floor isn’t a glitch or grade inflation. It reflects two design choices that business readers will recognize from real management.

Partial progress counts. A manager who correctly diagnoses a crisis but never closes the fix has still produced value — just not all of it. The do-nothing baseline still ends up owning correct assessments, avoided mistakes, and groundwork that a successor could build on. Scoring it zero would imply that diagnosis, triage, and preparation are worthless until the final signature lands. Anyone who has run a sales pipeline knows better.

A single breach of trust caps everything. The benchmark’s hardest rule, in its own words: “no amount of good work outweighs a breach of trust.” An agent could handle every crisis flawlessly, close every deal, and still see its total grade capped the moment it deceives a customer, hides a decision, or bends a rule. It’s the corporate equivalent of firing your top seller for cooking the books — competence doesn’t launder integrity.

Amazon

business AI compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Result Nobody Expected: Four Diagnoses, Two Signatures

Here’s what makes the floor matter. Every model in the field spotted every crisis and refused every manipulation attempt. Only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The benchmark’s summary of the gap: “Same diagnosis, same pitch — no signature.”

The decisive detail was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying AI in ecommerce or marketing: the agent that reads your internal documentation first beats the one that improvises from the customer’s words alone.

Amazon

AI trust and security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Pressure Test: Fake CEO, Patient Escalation, a Reporter’s Trap

The week included social engineering: fake CEO messages escalating over three stages, capped by a reporter’s disarmingly casual “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you want in anything with write access to real business systems.

Amazon

AI documentation and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Winning

The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Brilliance without follow-through is a familiar story to anyone who has managed talented people.

One fairness note worth flagging: K3 ran at the API’s default effort setting while the others ran at xhigh — and still took second at 93.

Why No 100s — and Why You Can Watch It Live

The distrust of round 100s is deliberate. A perfect score would mean a perfect week: every crisis handled, every file read, every deal closed, zero process slips. None of the frontier models got there. The final standings — gpt-5.6-sol 95, Kimi K3 93, Sonnet 5 88, Fable 5 77, Opus 4.8 73 — reward finishing, not flattering.

And this isn’t a one-off paper. Firmulate runs a live synthetic company with 13 employees, real money mechanics — €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for Business Buyers

Most AI demos measure how well a model chats. This one measures how well it manages — and its scoring philosophy is a checklist worth stealing for any AI procurement process:

  • Credit diagnosis, reward delivery. Partial progress has value, but the deal isn’t closed until it’s signed — the 26-point floor and the missing 100s exist for the same reason.
  • Make trust non-negotiable. One breach caps the total, no matter the scoreboard.
  • Test with your own files. The €55k deal was won by whoever read the company’s own documents, not whoever talked best.

The full league table and plain-language findings are at firmulate.com/benchmarks.html. If an AI agent is going to touch your revenue, that’s the kind of exam you want it to have passed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build Better Internal Reporting for Payment Teams

Discover how to enhance internal payment reports with real-time data, automation, and targeted insights to stay ahead of emerging threats and improve decision-making.

Understanding ACH Return Codes Simplified

Start unraveling the mysteries of ACH return codes with this simplified guide, essential for resolving payment discrepancies and optimizing transaction management.

Remote PCI Audits: Preparing Your Team and Tech Stack for a Virtual QSA

Keeping your team organized and tech-ready is crucial—discover key strategies to ensure a smooth remote PCI audit with a virtual QSA.

Why Back-Office Reconciliation Deserves More Attention

What makes back-office reconciliation crucial for your business’s accuracy and compliance? Discover how proper attention can protect your financial integrity.