🔍 Read the full analysis: Why Some AI Managers Always End With 26 In Benchmark Tests on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Benchmark tests for AI management tools show a consistent floor score of 26 and a ceiling near 95, highlighting how partial work is valued but trust breaches cap performance. The results reveal insights into AI decision-making under stress, as explored in this analysis.
In a recent benchmark league testing AI management tools under simulated business crises, the lowest score awarded was 26, even for models that did nothing, while the top scorer achieved 95. This scoring system emphasizes that partial progress is recognized but trust breaches sharply limit overall performance, a design choice that reveals how AI tools are evaluated in real-world management scenarios, as detailed in the original analysis.
The benchmark, conducted by Firmulate, involved four frontier AI models managing a small software company during a simulated week of crises, including customer issues, trust attacks, and operational failures. For more context, see the detailed report. The models’ scores ranged from 73 to 95, with the lowest baseline, labeled as the do-nothing scenario, receiving 26 points. This score reflects minimal but tangible effort, such as triaging emails or monitoring issues, acknowledging that even small actions have value in management.
The scoring system is intentionally designed to prevent inflation: a perfect score of 100 is considered suspicious, as it would suggest unmeasured or overly idealized performance. Instead, the maximum near 95 indicates models that successfully read documentation, refused manipulation attempts, and closed deals at full price. The benchmark also emphasizes trust: a single breach, such as failing to escalate or verify information, caps the total score, regardless of other achievements. This approach underscores that integrity in management is non-negotiable, even if operational performance is high.
Interestingly, models that read their own documentation and verified information were more successful at closing deals, highlighting the importance of thoroughness and accuracy. Conversely, models that failed to follow through or neglected to escalate issues scored lower, despite high rule-based analysis. The results demonstrate that discipline and follow-through are distinct skills, critical for trustworthy AI management. The benchmark also included social engineering tests, where models refused fake CEO messages and manipulative requests, indicating better handling of trust attacks than operational shortcuts.
Why Some AI Managers Always End With 26 In Benchmark Tests
Four frontier AI models ran a small software company through a simulated week of crises — customer failures, trust attacks, and operational breakdowns. The scoring floor was 26. The ceiling: 95. Here is what those two numbers reveal about trust in AI management.
A Week of Crises, Fully Auditable
The Firmulate benchmark league tests AI management models in realistic, high-pressure scenarios — not language fluency. Each model managed the same business through seven days of crises, scored on effectiveness, trustworthiness, and thoroughness.
Customer Support Failures
Inbound customer issues demanded triage, prioritization, and follow-through — small actions like monitoring and escalating emails earned measurable credit.
Trust Attacks
Social engineering tests included fake CEO messages and manipulative requests. Models that refused manipulation scored far better than those cutting operational corners.
Operational Decisions
Models that read their own documentation and verified information closed deals at full price — thoroughness proved a distinct skill from rule-based analysis.
From Do-Nothing Floor to Near-Perfect Ceiling
Why 26 and Why Not 100
The scoring system is intentionally designed to prevent inflation. Partial progress is recognized; trust breaches are decisive. Integrity in management is non-negotiable, even when operational performance is high.
“The cap at 95 and the floor at 26 are deliberate choices to prevent inflation and to recognize the value of partial progress while emphasizing trust as non-negotiable.”
— Thorsten MeyerThe Path to a Trustworthy 95
📄 Read Documentation
Models that studied their own docs and verified information closed deals at full price.
🛡️ Resist Manipulation
Refusing fake CEO messages and social engineering protected the trust score.
⚠️ Escalate Issues
Failing to escalate a single issue caps the total score — regardless of other wins.
✅ Close Deals Fully
Follow-through at full price with verification is what separates 73 from 95.
What Separates Top Scorers from the Rest
| Management Behavior | Top Scorer (≈95) | Mid Range (73–88) | Do-Nothing (26) |
|---|---|---|---|
| Read own documentation | ✓ Always | ~ Sometimes | ✗ Never |
| Refused fake CEO messages | ✓ Refused | ✓ Refused | ~ Ignored |
| Escalated critical issues | ✓ Consistently | ~ Inconsistent | ✗ Failed |
| Closed deals at full price | ✓ Full price | ~ Discounted | ✗ No deals |
| Minimal triage / monitoring | ✓ Exceeded | ✓ Exceeded | ~ Minimal only |
What This Means for AI Management
For organizations deploying AI managers, integrity and follow-through outrank raw operational capability. The cap discourages overestimating AI competence and encourages transparency about limitations — but open questions remain about how scores translate to live environments.
Q1Why does the benchmark cap scores at 95 and not 100?
A perfect score of 100 is seen as suspicious — a sign of unmeasured or overly idealized performance. Scores near 95 reflect top models that handle crises well but still have limitations.
Q2What does the score of 26 represent for the do-nothing baseline?
It reflects minimal but tangible effort — triaging emails, monitoring issues. Even small actions in management have value, so the floor prevents zero scores for partial progress.
Q3How important is trust in AI management according to this benchmark?
Central. A single trust breach — failing to escalate or verify — caps the total score regardless of operational success. Integrity is non-negotiable.
Q4Are these benchmark results applicable to real-world business operations?
While the crises are realistic, it remains uncertain how scores translate to live environments. Further validation is needed to confirm their predictive value in actual business settings.
Implications of Scoring Caps for AI Management Tools
The scoring system’s design, with a floor at 26 and a ceiling near 95, emphasizes that partial work is valued but trust breaches are decisive. For organizations deploying AI managers, this highlights the importance of integrity and follow-through over mere operational capability. The results suggest that AI tools must prioritize reading documentation, verifying information, and resisting manipulation to achieve trustworthy performance. The cap on the maximum score discourages overestimating AI competence and encourages transparency about limitations. Ultimately, this benchmark signals a shift toward valuing trustworthiness and consistent follow-through in AI management systems, which could influence future development and evaluation standards.
As an affiliate, we earn on qualifying purchases.
Background of the Firmulate Benchmark League
The benchmark league was launched by Firmulate to assess AI management models in realistic, high-pressure scenarios. Unlike traditional benchmarks that measure language fluency or problem-solving, this league evaluates how well AI tools manage a company during crises, including customer support, trust attacks, and operational decisions. The test involved four models managing the same business over seven days, with decisions fully auditable and scored based on effectiveness, trustworthiness, and thoroughness. The approach reflects growing industry interest in deploying AI for operational management, where trust and follow-through are critical.
The results, published in July 2026, follow a series of tests where models demonstrated varying levels of competence, with scores capped at 95 and a baseline of 26 for minimal effort. The league’s design intentionally avoids awarding perfect scores to prevent inflation, highlighting that partial work and trust are central to real-world AI management. This context underscores a broader industry shift toward evaluating AI systems not just on capabilities but on integrity and reliability.
“The cap at 95 and the floor at 26 are deliberate choices to prevent inflation and to recognize the value of partial progress while emphasizing trust as non-negotiable.”
— Thorsten Meyer
AI decision-making stress test software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Benchmark Limitations
It is not yet clear how these scores translate to real-world deployment beyond simulated crises. The benchmark’s focus on specific scenarios and trust breaches may not fully capture all aspects of operational AI management, such as long-term consistency or adaptability. Additionally, the impact of different model configurations, like the absence of effort parameters in some cases, remains under investigation. The industry awaits further validation on whether these scores accurately predict AI performance in live environments, and whether the trust caps will influence future AI development standards.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Following these results, industry stakeholders are likely to scrutinize how trust and follow-through are integrated into AI management tools. Future benchmarks may expand to include live operational environments, testing models over longer periods and more diverse scenarios. Developers might also focus on improving models’ ability to read documentation, verify information, and resist manipulation, aligning with the insights from this league. Additionally, the community may debate the appropriateness of the scoring caps and explore alternative evaluation metrics that balance operational efficiency with trustworthiness.
Meanwhile, organizations deploying AI managers are encouraged to consider these findings when selecting tools, prioritizing those that demonstrate integrity and thoroughness under pressure. The ongoing evolution of benchmarking standards will shape how AI management systems are built and trusted in critical business functions.
AI operational monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark cap scores at 95 and not 100?
The benchmark designers see a perfect score of 100 as suspicious, indicating unmeasured or overly idealized performance. Scores near 95 are considered realistic for top models that handle crises well but still have limitations.
What does the score of 26 represent for the do-nothing baseline?
The score of 26 reflects minimal effort, such as triaging emails or monitoring issues, recognizing that even small actions in management have value. It serves as a floor to prevent zero scores for partial progress.
How important is trust in AI management according to this benchmark?
Trust is central; a single breach caps the total score regardless of operational success. The benchmark emphasizes that integrity and follow-through are non-negotiable in AI management tools.
Will these results influence future AI development?
Yes, the results highlight the importance of reading documentation, verifying information, and resisting manipulation, which are likely to become key focus areas for developers aiming for trustworthy AI management systems.
Are these benchmark results applicable to real-world business operations?
While the benchmark simulates realistic crises, it remains uncertain how scores translate to live environments. Further validation is needed to confirm their predictive value in actual business settings.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
