📊 Full opportunity report: Can A Management Test Reveal The Authentic Work Style Of AI? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A new experiment tests AI models in managing a simulated company under crisis. Results suggest management tests can reveal distinct AI work styles, but challenges remain in assessing operational effectiveness.
A management experiment conducted by Firmulate has demonstrated that AI models can be evaluated based on their decision-making styles in realistic business scenarios. The test involves AI models managing a simulated company through a week of crises, revealing differences in diligence, discipline, and follow-through. This approach offers a new way to assess AI’s management capabilities beyond traditional performance metrics, similar to the insights from the original analysis.
The experiment, hosted on firmulate.com, involved five frontier AI models tasked with running a small software company facing a series of crises, including customer issues, financial pressures, and security threats. Each model was given identical problems, and their decisions were observed and scored based on effectiveness, trustworthiness, and completion of critical tasks.
The results, published in July 2026, ranked the models from highest to lowest score: GPT-5.6-SOL leading with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The experiment highlighted that models could recognize crises and refuse manipulative requests, but their ability to complete decisive actions varied significantly.
For example, Opus 4.8 produced detailed analyses but often failed to follow through on operational steps, such as closing deals or escalating issues appropriately. Conversely, models like Kimi K3 and GPT-5.6-SOL demonstrated stronger operational discipline, completing critical tasks and securing business outcomes.
Can a Management Test Reveal the Authentic Work Style of AI?
A simulated week of customer emergencies, financial pressure and security threats suggests that management tests can expose distinct AI work styles—especially the gap between diagnosing a crisis and actually resolving it.
Identical business scenarios revealed meaningful differences in discipline, diligence and follow-through.
Models that reasoned at length did not always complete the operational actions needed to protect the company.
Real-world reliability remains insufficiently tested for autonomous management.
Management style emerges under pressure
The Firmulate experiment gave each model the same company, the same crises and the same opportunity to act. Differences appeared not only in what models understood, but in what they completed.
Can it spot the crisis?
Models were tested on identifying customer, financial and security risks before those risks compounded.
Can it choose a sound response?
The simulation exposed how models prioritize trade-offs, resist manipulation and judge when escalation is necessary.
Can it finish the job?
The decisive separator was follow-through: closing deals, escalating issues and completing critical operational steps.
Five models, five operational profiles
Scores combined effectiveness, trustworthiness and task completion. The 22-point spread indicates that management simulations can distinguish behavioral patterns hidden by conventional benchmarks.
| Rank | Model | Score | Operational discipline | Follow-through | Observed profile |
|---|---|---|---|---|---|
| 01 | GPT-5.6-SOL | 95 | ✓ Strong | ✓ Consistent | Decisive, reliable execution |
| 02 | Kimi K3 | 93 | ✓ Strong | ✓ Consistent | Secured critical outcomes |
| 03 | Sonnet 5 | 88 | ~ Capable | ~ Variable | Balanced, with some execution gaps |
| 04 | Fable 5 | 77 | ~ Uneven | ✗ Limited | Recognized issues, missed key actions |
| 05 | Opus 4.8 | 73 | ~ Analytical | ✗ Incomplete | Detailed analysis without decisive closure |
The execution gap is measurable
All five systems could engage with complex crises. Their ability to translate reasoning into completed business actions varied substantially.
“Thorough analysis alone isn’t enough; effective management requires AI to act decisively and reliably under pressure.”AI researcher involved in the experiment
From crisis signal to business outcome
A credible management test must trace the entire decision chain. Recognizing a threat earns little if the final action is delayed, incomplete or never verified.
Detect
Identify the customer, cash-flow or security threat.
Prioritize
Judge urgency, impact and competing demands.
Decide
Select a trustworthy and commercially sound response.
Execute
Complete the operational steps and escalate when needed.
Verify
Confirm resolution and protect the intended outcome.
Promising evidence, not a final verdict
The simulation advances AI evaluation beyond accuracy and speed, but a controlled week cannot establish long-term reliability in an unpredictable organization.
Evaluate behavior before deployment
Operational AI should be assessed in realistic scenarios that test judgment and action together.
- Measure completion, not just recommendation quality.
- Test resistance to manipulation and unsafe requests.
- Record escalation choices and unresolved tasks.
- Compare reliability across repeated crisis scenarios.
Simulation is not the real world
The experiment leaves several consequential questions unanswered.
- Will observed work styles persist over months?
- How will models respond to novel, ambiguous events?
- How much can fine-tuning alter management behavior?
- Do scores transfer across industries and company sizes?
What decision-makers should ask
Management-style testing is most useful as a pre-deployment risk lens—not as proof that AI can replace accountable human leadership.
Can AI replace human managers?
Not on this evidence. Models can support crisis analysis and decision-making, but judgment, accountability and reliable execution still require human oversight.
How is this different from a benchmark?
Traditional benchmarks score isolated outputs. Management tests observe connected decisions, execution steps and outcomes across an evolving situation.
Are the models deployment-ready?
They are better viewed as decision-support systems until reliability is validated across repeated, unpredictable and higher-stakes environments.
What should companies measure?
Track effectiveness, trustworthiness, escalation judgment, decisive action, completion rates and whether the claimed resolution was actually verified.
Implications for AI Management Evaluation
This experiment demonstrates that management-style tests can differentiate AI models based on their decision-making behaviors in realistic scenarios. It suggests that enterprises evaluating AI for operational roles should consider not just analytical accuracy but also the model’s ability to execute and follow through on decisions, which are crucial for real-world management.
While promising, the findings also reveal limitations—more analysis does not automatically translate into better management. Operational discipline and trustworthiness are key factors that influence AI’s practical utility in business settings.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Testing Approaches
Traditional AI evaluations focus on accuracy, speed, or task-specific performance, often in controlled environments. Recent developments aim to assess AI in more complex, real-world contexts, including management simulations that mimic business crises. Firmulate’s live experiment is among the first to test multiple models in a simulated operational environment, providing insights into their management styles and decision-making traits.
Previous efforts have highlighted AI’s strengths in analysis but less so in execution. This experiment bridges that gap by observing how models handle both diagnosis and decisive action, emphasizing the importance of operational discipline in AI management.
“Testing AI models in realistic management scenarios reveals critical differences in their ability to follow through and execute, which are often overlooked in traditional benchmarks.”
— Firmulate representative
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Test Outcomes
It remains unclear how these results will generalize to real-world business environments outside the controlled simulation. The long-term reliability of these models in operational roles, especially under unpredictable conditions, has yet to be tested. Additionally, the impact of different training methods or fine-tuning on management behaviors is still under investigation.
AI decision-making assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Evaluation
Researchers plan to expand the experiment by incorporating more diverse scenarios and additional AI models. Enterprises may soon have access to similar testing frameworks, allowing them to evaluate AI management personalities before deployment. Further studies will also explore how to improve models’ operational discipline and trustworthiness in complex tasks.
AI operational discipline evaluation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Can AI models truly replace human managers?
Current experiments suggest AI can support management tasks by analyzing crises and making decisions, but full replacement of human managers remains uncertain due to limitations in operational discipline and judgment.
What makes management-style testing different from traditional AI benchmarks?
Management tests evaluate AI’s ability to make decisions, execute actions, and handle crises in realistic scenarios, focusing on operational discipline rather than just analytical performance.
Are these AI models ready for real-world business management?
While promising, these models still require further validation in unpredictable environments. They are best viewed as decision-support tools at this stage.
How can companies evaluate AI for operational roles?
Companies should consider management-style simulations that test AI decision-making under pressure, including follow-through and trustworthiness, before deployment.
Source: ThorstenMeyerAI.com