AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Agents Need Practice With Business Problems Before Going Live on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s live simulation, the Crucible League, ran five frontier AI models through a software company’s worst week in July 2026. Every model detected each crisis and refused manipulation attempts, but only two closed a €55,000 deal their own analysis justified, showing why agents need company-specific practice before touching live operations.

The final Crucible League, a live simulation completed in July 2026, tested five frontier AI models on running the same small software company through its most difficult week — and found that while every model detected every crisis and refused every manipulation attempt, only two signed a €55,000 deal that their own analysis had earned. The experiment, as detailed in the original analysis, published on firmulate.com, is being positioned as evidence that AI agents need rehearsal against realistic business problems, ideally a company’s own data, before being deployed near live operations.

Frontier models each ran the identical simulated company through a deliberately bad week, with every decision versioned and auditable. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted toward scores, but a single breach of trust capped the total — the experiment’s stated rule was that “no amount of good work outweighs a breach of trust.”

The headline result was not missed emergencies. According to the published results, all models spotted every crisis and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request framed as “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The decisive gap appeared after diagnosis. The winning models found a competitor weakness buried two document references deep in the company’s own files — not in the customer event itself — and used it to close the deal at full price, worth +€4,583 in monthly recurring revenue. The experiment’s summary of the failure mode: “Same diagnosis, same pitch — no signature.” Opus 4.8 illustrated a related finding: it was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last after leaving the close on the table and attempting to write into a locked department instead of escalating.

At a glance
reportWhen: final league results completed July 202…
The developmentFirmulate published final results of its July 2026 Crucible League, in which frontier AI models ran a simulated software company through a crisis week, and is extending the exercise into read-only enterprise pilots.
AI Agents Need Practice With Business Problems Before Going Live
The Crucible League · July 2026 · Firmulate

AI Agents Need Practice With Business Problems Before Going Live

Five frontier models ran a software company through its worst week in a live simulation. All spotted every crisis and refused every manipulation — but only two closed a €55,000 deal their own analysis justified. The gap between diagnosis and action is why agents need rehearsal against realistic, company-specific scenarios before deployment.

“No amount of good work outweighs a breach of trust.”
Firmulate Experiment Rules
Models Tested5
Crises Missed0
Deals Closed2/5
Write-Backs0
Top Score · gpt-5.6-sol95
Do-Nothing Baseline26
Self-Learned Playbook Rules680+
Auditable Decisions in Quiz242
01 · Final Standings

Same Week, Same Company, Very Different Outcomes

Each frontier model managed the identical simulated software company through a deliberately bad week. Partial progress counted toward scores — but a single breach of trust capped the total.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing
26
Caveat

Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh. Firmulate states the standings are a record of this single experiment — the configuration difference is context, not a controlled variable.

02 · The Gap

Why Diagnosis Alone Doesn’t Qualify an Agent

Every model detected every crisis and refused every manipulation attempt — including staged fake CEO messages escalating over three stages and a reporter’s request framed as “just one yes/no, on background.” All five refused. The decisive gap appeared after diagnosis.

Capability · Detection

Spotting the Crisis

All five models identified every emergency and treated manipulation attempts as suspected approval-bypasses or impersonation. A polished demo proves this layer — and only this layer.

Capability · Evidence

Finding Buried Intel

The winners uncovered a competitor weakness two document references deep in the company’s own files — not in the customer event — and used it to close the €55,000 deal at full price: +€4,583 in monthly recurring revenue.

Capability · Boundaries

Respecting Locked Doors

A weaker version of the boundary-discipline failure appeared in four of five models: attempting to write into a locked department instead of escalating. Thoroughness alone didn’t save Opus 4.8.

Opus 4.8

The most thorough participant — 80 learned rules added, deepest analyses produced — finished last after leaving the close on the table. “Same diagnosis, same pitch — no signature.”

03 · How It Works

From Public Simulation to Your Own Data

Firmulate is extending the live experiment into read-only enterprise pilots. Nothing writes back to real systems.

1

Live Simulation

A synthetic software company with 13 synthetic employees, €105,000/month burn against €2,300 MRR, and a public cash countdown.

2

Versioned Decisions

Every workday and decision is versioned and auditable — 242 real, unedited management decisions feed a public guessing quiz.

3

Enterprise Pilot

The same wargame runs against a read-only export of a company’s own data, testing crisis scenarios safely.

4

Board Report

Each pilot produces model rankings plus weak points in the company’s existing playbooks — a pre-deployment benchmark.

04 · Head-to-Head

What Each Model Did With the Worst Week

Model Score Detected Crises Refused Manipulation Closed €55K Deal Notable Behavior
gpt-5.6-sol95 ✓ All✓ All attempts✓ Full price Found competitor weakness deep in company files
Kimi K393 ✓ All✓ All attempts✓ Full price On-record: “suspected approval-bypass / possible impersonation”
Sonnet 588 ✓ All✓ All attempts✗ Left on table Partial boundary-discipline weakness
Fable 577 ✓ All✓ All attempts✗ Left on table Partial boundary-discipline weakness
Opus 4.873 ✓ All✓ All attempts✗ Left on table 80 learned rules, deepest analyses — wrote into locked dept instead of escalating
05 · On the Record

Voices From the Crucible

“No amount of good work outweighs a breach of trust.”

Firmulate Experiment Rules

“Same diagnosis, same pitch — no signature.”

Firmulate Published Results

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 · On-Record Reasoning
06 · Key Questions & Limits

What Buyers Should Ask Before Deployment

What is the Crucible League?

A live simulation on firmulate.com in which frontier AI models each managed the same small synthetic software company through its worst week. The final league completed in July 2026, with every decision versioned and auditable.

Which models scored highest?

gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73 — against a do-nothing baseline of 26. Kimi K3 ran without an effort parameter; the others ran at xhigh.

Did any model fall for manipulation attempts?

No. All five refused every attempt, including staged fake CEO messages and a reporter’s on-background confirmation request.

Is company data at risk in the pilot?

The pilot runs a wargame against a read-only export of a company’s own data. Firmulate states nothing writes back to real systems during the exercise.

Why did the most thorough model finish last?

Opus 4.8 added 80 learned rules and produced the deepest analyses, but left the €55,000 deal unclosed and wrote into a locked department instead of escalating.

What are the limits of the rankings?

One synthetic company, one difficult week. Transfer to other industries, larger organizations, or longer horizons is not established — and no independent pilot results have been published yet.

Why Diagnosis Alone Doesn’t Qualify an Agent

The results point to a practical gap for companies adopting AI automation: an agent can recognize a situation correctly and make a persuasive case, yet still fail to act on information already available inside the business. For buyers evaluating AI tools, a polished demo shows what an agent says, but not whether it will finish the job when a real business is under pressure.

The league suggests that spotting a crisis and refusing a scam are not the whole job. Models also need to find relevant evidence, close justified opportunities, and respect boundaries when their first route is blocked. A weaker version of that boundary-discipline weakness appeared in four of the five models, according to the published findings. That makes a case for testing agents against realistic, company-specific scenarios before deployment rather than relying on vendor demonstrations.

How the Simulation and Enterprise Pilot Work

Firmulate’s live experiment runs a synthetic software company with 13 synthetic employees and real money mechanics: a burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays. Readers can follow the simulation live and take a quiz built from 242 real, unedited management decisions, guessing which model made each choice.

The enterprise offering extends the concept. A pilot runs the same style of wargame against a read-only export of a company’s own data, testing crisis scenarios and producing a board report with model rankings and weak points in the company’s existing playbooks. Firmulate states that nothing writes back to real systems. Companies interested in a pilot are directed to the pilot page or contact@firmulate.com.

“No amount of good work outweighs a breach of trust.”

— Firmulate experiment rules

Caveats in the Model Rankings

The experiment itself flags a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate states the standings are a record of this single experiment, with that configuration difference part of the context rather than a controlled variable.

Other limits remain. The simulation covers one synthetic company over one difficult week; how the results transfer to other industries, larger organizations, or longer time horizons is not established. The enterprise pilot reports are described but no independent customer results have been published, and the benchmark scores are specific to this scenario rather than a general capability measure.

From Watching a Synthetic Company to Testing Your Own

Firmulate is moving from the public experiment toward enterprise pilots using read-only data exports, each producing a board report with model rankings and identified playbook weaknesses. The live simulation and full benchmark results remain available at firmulate.com/live and firmulate.com/benchmarks.html, so subsequent runs or updated model versions could be compared against the July 2026 baseline. Whether companies adopt wargame-style testing as a standard pre-deployment step for AI agents — the outcome the experiment implicitly argues for — will depend on results from those first pilots, which have not yet been published.

Key Questions

What is the Crucible League?

It is a live simulation run on firmulate.com in which frontier AI models each managed the same small synthetic software company through its worst week. The final league was completed in July 2026, with every decision versioned and auditable.

Which models scored highest?

gpt-5.6-sol finished first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate notes Kimi K3 ran without an effort parameter while the others ran at xhigh, so the comparison carries that caveat.

Did any model fall for manipulation attempts?

No. According to the published results, all five models refused every manipulation attempt, including staged fake CEO messages and a reporter’s request for an on-background confirmation.

What is the enterprise pilot and is company data at risk?

The pilot runs a wargame against a read-only export of a company’s own data and produces a board report with model rankings and playbook weak points. Firmulate states that nothing writes back to real systems during the exercise.

Why did the most thorough model finish last?

Opus 4.8 added 80 learned rules and produced the deepest analyses, but left the €55,000 deal unclosed and attempted to write into a locked department instead of escalating. The experiment treats this as evidence that thoroughness does not guarantee effective action.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bird Construction Inc. Announces Release Date And Conference Call For 2026 Second Quarter Financial Results

Bird Construction has announced the timing for its Q2 2026 results release and conference call, with financial figures still pending.

Is Walmart Open On 4Th Of July

Walmart is open on July 4th, with store hours varying by location. Find out what to expect for Independence Day shopping.

Inuvo To Host Second Quarter 2026 Financial Results Conference Call On Tuesday, August 11Th At 4:15 P.M. ET

Inuvo announced it will host its second quarter 2026 financial results conference call on Tuesday, August 11th at 4:15 p.m. ET, providing updates to investors.

World War II Fighter Wreck Of America’s Top Ace Recovered From Jungles Of Papua New Guinea

A World War II fighter aircraft belonging to America’s top ace has been recovered from the jungles of Papua New Guinea, confirming long-held suspicions of its location.