AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Can A Management Test Reveal The Authentic Work Style Of AI? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A new experiment tests AI models in managing a simulated company under crisis. Results suggest management tests can reveal distinct AI work styles, but challenges remain in assessing operational effectiveness.

A management experiment conducted by Firmulate has demonstrated that AI models can be evaluated based on their decision-making styles in realistic business scenarios. The test involves AI models managing a simulated company through a week of crises, revealing differences in diligence, discipline, and follow-through. This approach offers a new way to assess AI’s management capabilities beyond traditional performance metrics, similar to the insights from the original analysis.

The experiment, hosted on firmulate.com, involved five frontier AI models tasked with running a small software company facing a series of crises, including customer issues, financial pressures, and security threats. Each model was given identical problems, and their decisions were observed and scored based on effectiveness, trustworthiness, and completion of critical tasks.

The results, published in July 2026, ranked the models from highest to lowest score: GPT-5.6-SOL leading with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The experiment highlighted that models could recognize crises and refuse manipulative requests, but their ability to complete decisive actions varied significantly.

For example, Opus 4.8 produced detailed analyses but often failed to follow through on operational steps, such as closing deals or escalating issues appropriately. Conversely, models like Kimi K3 and GPT-5.6-SOL demonstrated stronger operational discipline, completing critical tasks and securing business outcomes.

At a glance
reportWhen: developing; results announced July 2026
The developmentAn ongoing experiment evaluates if management tests can expose authentic AI work styles through real business simulations.
Can a Management Test Reveal the Authentic Work Style of AI?
AI Management Evaluation · July 2026

Can a Management Test Reveal the Authentic Work Style of AI?

A simulated week of customer emergencies, financial pressure and security threats suggests that management tests can expose distinct AI work styles—especially the gap between diagnosing a crisis and actually resolving it.

Short answer Yes—with limits.

Identical business scenarios revealed meaningful differences in discipline, diligence and follow-through.

Core finding Analysis is not execution.

Models that reasoned at length did not always complete the operational actions needed to protect the company.

Deployment view Decision support first.

Real-world reliability remains insufficiently tested for autonomous management.

5 Frontier models
1 week Simulated operations
95 Top score
22 pts Performance spread
3 Scoring dimensions

Management style emerges under pressure

The Firmulate experiment gave each model the same company, the same crises and the same opportunity to act. Differences appeared not only in what models understood, but in what they completed.

Recognition

Can it spot the crisis?

Models were tested on identifying customer, financial and security risks before those risks compounded.

Judgment

Can it choose a sound response?

The simulation exposed how models prioritize trade-offs, resist manipulation and judge when escalation is necessary.

Execution

Can it finish the job?

The decisive separator was follow-through: closing deals, escalating issues and completing critical operational steps.

Five models, five operational profiles

Scores combined effectiveness, trustworthiness and task completion. The 22-point spread indicates that management simulations can distinguish behavioral patterns hidden by conventional benchmarks.

Rank Model Score Operational discipline Follow-through Observed profile
01 GPT-5.6-SOL 95 ✓ Strong ✓ Consistent Decisive, reliable execution
02 Kimi K3 93 ✓ Strong ✓ Consistent Secured critical outcomes
03 Sonnet 5 88 ~ Capable ~ Variable Balanced, with some execution gaps
04 Fable 5 77 ~ Uneven ✗ Limited Recognized issues, missed key actions
05 Opus 4.8 73 ~ Analytical ✗ Incomplete Detailed analysis without decisive closure
Reported July 2026 · Scores reflect the described Firmulate simulation, not validated real-world management performance.

The execution gap is measurable

All five systems could engage with complex crises. Their ability to translate reasoning into completed business actions varied substantially.

GPT-5.6-SOL
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
“Thorough analysis alone isn’t enough; effective management requires AI to act decisively and reliably under pressure.”
AI researcher involved in the experiment

From crisis signal to business outcome

A credible management test must trace the entire decision chain. Recognizing a threat earns little if the final action is delayed, incomplete or never verified.

01

Detect

Identify the customer, cash-flow or security threat.

02

Prioritize

Judge urgency, impact and competing demands.

03

Decide

Select a trustworthy and commercially sound response.

04

Execute

Complete the operational steps and escalate when needed.

05

Verify

Confirm resolution and protect the intended outcome.

Promising evidence, not a final verdict

The simulation advances AI evaluation beyond accuracy and speed, but a controlled week cannot establish long-term reliability in an unpredictable organization.

What enterprises can use now

Evaluate behavior before deployment

Operational AI should be assessed in realistic scenarios that test judgment and action together.

  • Measure completion, not just recommendation quality.
  • Test resistance to manipulation and unsafe requests.
  • Record escalation choices and unresolved tasks.
  • Compare reliability across repeated crisis scenarios.
What remains unclear

Simulation is not the real world

The experiment leaves several consequential questions unanswered.

  • Will observed work styles persist over months?
  • How will models respond to novel, ambiguous events?
  • How much can fine-tuning alter management behavior?
  • Do scores transfer across industries and company sizes?

What decision-makers should ask

Management-style testing is most useful as a pre-deployment risk lens—not as proof that AI can replace accountable human leadership.

Can AI replace human managers?

Not on this evidence. Models can support crisis analysis and decision-making, but judgment, accountability and reliable execution still require human oversight.

How is this different from a benchmark?

Traditional benchmarks score isolated outputs. Management tests observe connected decisions, execution steps and outcomes across an evolving situation.

Are the models deployment-ready?

They are better viewed as decision-support systems until reliability is validated across repeated, unpredictable and higher-stakes environments.

What should companies measure?

Track effectiveness, trustworthiness, escalation judgment, decisive action, completion rates and whether the claimed resolution was actually verified.

Bottom line Management tests can reveal authentic differences in AI work style—but operational discipline, not analytical fluency, is the capability that matters most.

Source: ThorstenMeyerAI.com · Experiment status: developing · Results reported July 2026

Powered by Thorsten Meyer AI
AI Management Brief

Implications for AI Management Evaluation

This experiment demonstrates that management-style tests can differentiate AI models based on their decision-making behaviors in realistic scenarios. It suggests that enterprises evaluating AI for operational roles should consider not just analytical accuracy but also the model’s ability to execute and follow through on decisions, which are crucial for real-world management.

While promising, the findings also reveal limitations—more analysis does not automatically translate into better management. Operational discipline and trustworthiness are key factors that influence AI’s practical utility in business settings.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Testing Approaches

Traditional AI evaluations focus on accuracy, speed, or task-specific performance, often in controlled environments. Recent developments aim to assess AI in more complex, real-world contexts, including management simulations that mimic business crises. Firmulate’s live experiment is among the first to test multiple models in a simulated operational environment, providing insights into their management styles and decision-making traits.

Previous efforts have highlighted AI’s strengths in analysis but less so in execution. This experiment bridges that gap by observing how models handle both diagnosis and decisive action, emphasizing the importance of operational discipline in AI management.

“Testing AI models in realistic management scenarios reveals critical differences in their ability to follow through and execute, which are often overlooked in traditional benchmarks.”

— Firmulate representative

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Test Outcomes

It remains unclear how these results will generalize to real-world business environments outside the controlled simulation. The long-term reliability of these models in operational roles, especially under unpredictable conditions, has yet to be tested. Additionally, the impact of different training methods or fine-tuning on management behaviors is still under investigation.

Amazon

AI decision-making assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Management Evaluation

Researchers plan to expand the experiment by incorporating more diverse scenarios and additional AI models. Enterprises may soon have access to similar testing frameworks, allowing them to evaluate AI management personalities before deployment. Further studies will also explore how to improve models’ operational discipline and trustworthiness in complex tasks.

Amazon

AI operational discipline evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Can AI models truly replace human managers?

Current experiments suggest AI can support management tasks by analyzing crises and making decisions, but full replacement of human managers remains uncertain due to limitations in operational discipline and judgment.

What makes management-style testing different from traditional AI benchmarks?

Management tests evaluate AI’s ability to make decisions, execute actions, and handle crises in realistic scenarios, focusing on operational discipline rather than just analytical performance.

Are these AI models ready for real-world business management?

While promising, these models still require further validation in unpredictable environments. They are best viewed as decision-support tools at this stage.

How can companies evaluate AI for operational roles?

Companies should consider management-style simulations that test AI decision-making under pressure, including follow-through and trustworthiness, before deployment.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX acquired AI coding startup Cursor for $60 billion in stock, a move that analysts see as a strategic bargain amid rapid growth and vertical integration.

The Forward-Deploy Pivot: Why Anthropic and OpenAI Are Becoming Consulting Firms in the Same Week

Anthropic and OpenAI are establishing enterprise services firms, signaling a strategic move into consulting and AI-driven outcomes, challenging traditional industry players.

How L3 Data Unlocks Deep B2B Payment Insights—If You Capture It Correctly

The key to unlocking deep B2B payment insights lies in capturing L3 data correctly, revealing details that can transform your analysis—find out how.