TL;DR
A live AI management test by Firmulate demonstrates that current models can identify crises but often fail to complete decisions, highlighting the need for new evaluation standards focused on management quality. The experiment underscores the importance of trust, execution, and consequence management in AI systems.
Firmulate has launched a live management benchmark experiment that tests AI models in a simulated company environment facing real crises during its worst week. The results, announced in July 2026, show that while models can identify issues and resist manipulation, they often fail to complete critical decisions or sign deals, exposing a gap in current AI evaluation methods. This matters because it shifts the focus from chat quality to management effectiveness, which is crucial for deploying AI in real business operations.
The experiment involved five AI managers competing in the Crucible League, with GPT-5.6-SOL achieving the highest score of 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The company simulated a small software firm with €2.3k MRR burning €105k monthly, facing crises like churn, PR issues, and sales negotiations. All models successfully diagnosed crises and resisted manipulation attempts, such as fake CEO messages and background approval requests.
However, significant gaps emerged in decision-making and trust. For instance, only two models signed a €55,000 deal after analysis, despite correctly diagnosing the opportunity. The failure was traced to models not retrieving critical facts from internal documents, which changed the commercial outcome. The experiment also highlighted that models could recognize manipulation but still faltered in completing managerial tasks, such as escalating issues or finalizing agreements.
Interestingly, the most detailed model, Opus 4.8, added extensive rules and analysis but finished last, illustrating that more effort and activity do not necessarily translate into better management outcomes. The benchmark emphasizes that effective management involves not just diagnosis but also execution, trust, and consequence management—areas where current AI models still struggle.
Implications for AI in Business Management
This experiment demonstrates that current AI models excel at identifying problems but often fall short in the critical areas of decision execution, trustworthiness, and consequence management. For organizations considering AI for operational roles, this highlights the importance of evaluating models not only on their diagnostic accuracy but also on their ability to complete tasks, maintain honesty, and handle real-world pressures. The findings suggest a need to develop new benchmarks that focus on management quality rather than just chat or coding performance, potentially reshaping how AI readiness is assessed in enterprise contexts.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background
Traditional AI benchmarks mainly assess technical output, such as coding accuracy or conversational quality, often neglecting how models perform in dynamic, high-stakes situations. The Firmulate experiment fills this gap by simulating a real company environment with ongoing crises, real money mechanics, and versioned decisions. Prior to this, AI evaluation relied heavily on static tests or isolated tasks, which do not reflect the complexities of managing a business under pressure. The experiment builds on recent recognition that AI systems need to handle not just information retrieval but also consequence-aware management.
Since its inception, the concept of benchmarking AI in operational settings has gained traction, but no standardized tests have yet integrated the full scope of management responsibilities. This experiment marks a significant step toward that goal, emphasizing that the true measure of AI readiness for enterprise use involves its capacity to manage trust, escalate appropriately, and complete decisions under pressure.
“The real test for AI in management isn’t just diagnosing problems; it’s completing the job without compromising trust or execution.”
— Thorsten Meyer, lead researcher
AI crisis management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Challenges in AI Management Evaluation
While the experiment provides valuable insights, several questions remain. It is not yet clear how well these findings generalize to larger organizations or different industries. The long-term reliability of models in maintaining trust and completing complex tasks under varied conditions is still uncertain. Additionally, the impact of integrating such benchmarks into enterprise procurement processes remains to be seen, as does the development of standardized management-focused evaluation metrics.
enterprise AI trust and execution solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Benchmark Development and Adoption
Researchers and industry leaders are expected to build on these findings by developing more comprehensive benchmarks that include management and consequence-based tasks. Companies considering AI for operational roles should start conducting internal wargames and scenario testing to assess models’ ability to handle real-world pressures. Regulatory and standards bodies may also begin proposing new evaluation criteria focused on trustworthiness, decision completeness, and escalation protocols. The ongoing refinement of these benchmarks will determine how quickly and effectively AI can be integrated into core management functions.
AI performance evaluation benchmarks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main purpose of the Firmulate experiment?
The experiment aims to evaluate AI models’ ability to manage a simulated company crisis, focusing on decision-making, trust, and execution, rather than just diagnostic accuracy.
Why are current benchmarks insufficient for management tasks?
Existing benchmarks primarily measure technical performance like coding or chat quality, neglecting critical management skills such as completing decisions, maintaining trust, and managing consequences under pressure.
What does the experiment reveal about AI safety and trust?
The models effectively resist manipulation attempts, such as fake messages, but often fail to follow through on decisions or escalate issues properly, exposing gaps in trust and execution.
How might this impact AI deployment in businesses?
Businesses may need to adopt new evaluation standards that test AI models’ management capabilities, ensuring they can handle real-world pressures without compromising trust or effectiveness.
What are the next steps for AI benchmarking?
Developing more comprehensive, management-focused benchmarks and conducting internal scenario tests will be key to advancing AI readiness for operational roles.
Source: ThorstenMeyerAI.com