AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A polished demo can hide the moment an AI agent fails to finish the job

For a business weighing AI for customer support, sales or operations, fluent answers are only part of the test. Can an agent spot the relevant detail in company records, make the right call under pressure and follow through? Firmulate’s latest company simulation puts those questions into a live, auditable contest—and its results complicate the idea that the leading choices are already settled.

Amazon

AI customer support chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A newcomer nearly tops the table

In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. K3 found a buried security concern, secured a €55,000 deal worth €4,583 in monthly recurring revenue, retained a customer at risk of leaving and resisted the manipulation attempts. It also had just one deviation, the fewest in the field.

The ranking is a close race at the top, not a clean sweep. But it shows that a less familiar model can perform alongside—and ahead of—several Western frontier models in a demanding business scenario. Choosing a model based on reputation alone now looks like a wager. The results and plain-language findings are available on Firmulate’s benchmark page.

Amazon

AI decision-making simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The test is about execution, not just diagnosis

Firmulate ran each model through the same bad week at a small software company, with the same customers, crises and temptations. Every decision was versioned and auditable. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price. That finding has a familiar business lesson: having a sound analysis is not the same as acting on it, and a crucial clue can sit outside the most obvious conversation or dashboard.

Trust under pressure, and discipline in practice

The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The stated principle was “no amount of good work outweighs a breach of trust.”

Opus 4.8 presents a different caution. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four other participants. Thoroughness, then, did not guarantee follow-through.

Amazon

AI trust and ethics training software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A live company, with a fairness caveat

Firmulate describes the experiment as a running company emulator, not a slide presentation. Its synthetic company has 13 employees, a public cash countdown and more than 680 self-learned playbook rules; each workday is versioned. The business mechanics include monthly burn of €105,000 against €2,300 in monthly recurring revenue. Readers can watch the company at Firmulate.

The comparison has an important qualification: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs alongside the scores when interpreting the standings. The experiment is useful evidence about these runs, not a universal verdict on how every model will behave in every company.

Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. For enterprises, its pilot applies the same wargame to a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI model performance testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work you need done

For business and marketing teams, the practical question is not simply which model writes the most convincing answer. It is whether an agent finds the needed evidence, protects trust and carries a decision through to completion. Firmulate’s league suggests those qualities can vary—and that reputation alone cannot settle the choice. Run a test against the work your own business needs done before handing an agent the keys.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Revolutionizing Transactions: the Power of Echecks

Open the door to modernized transactions with eChecks, providing enhanced efficiency and reduced costs – find out how!

Rackmount UPS Units: Overkill or Smart Resilience?

Probing whether rackmount UPS units are overkill or smart resilience reveals key features that could transform your equipment’s safety and longevity.

Network Racks and Cabinets Matter More in Retail Than You’d Expect

Beneath their simple appearance, network racks and cabinets are crucial for retail success, ensuring security, reliability, and optimal performance—find out why they matter more than you think.

Third‑Party Risk Management: Vetting Vendors Without Slowing Growth

Managing third-party risks without hindering growth requires strategic vetting; discover how to strike the right balance by continuing to read.