The AI Leaderboard That Matters Starts After The Demo Ends
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

A live AI management test by Firmulate demonstrates that current models can identify crises but often fail to complete decisions, highlighting the need for new evaluation standards focused on management quality. The experiment underscores the importance of trust, execution, and consequence management in AI systems.

Firmulate has launched a live management benchmark experiment that tests AI models in a simulated company environment facing real crises during its worst week. The results, announced in July 2026, show that while models can identify issues and resist manipulation, they often fail to complete critical decisions or sign deals, exposing a gap in current AI evaluation methods. This matters because it shifts the focus from chat quality to management effectiveness, which is crucial for deploying AI in real business operations.

The experiment involved five AI managers competing in the Crucible League, with GPT-5.6-SOL achieving the highest score of 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The company simulated a small software firm with €2.3k MRR burning €105k monthly, facing crises like churn, PR issues, and sales negotiations. All models successfully diagnosed crises and resisted manipulation attempts, such as fake CEO messages and background approval requests.

However, significant gaps emerged in decision-making and trust. For instance, only two models signed a €55,000 deal after analysis, despite correctly diagnosing the opportunity. The failure was traced to models not retrieving critical facts from internal documents, which changed the commercial outcome. The experiment also highlighted that models could recognize manipulation but still faltered in completing managerial tasks, such as escalating issues or finalizing agreements.

Interestingly, the most detailed model, Opus 4.8, added extensive rules and analysis but finished last, illustrating that more effort and activity do not necessarily translate into better management outcomes. The benchmark emphasizes that effective management involves not just diagnosis but also execution, trust, and consequence management—areas where current AI models still struggle.

At a glance
reportWhen: ongoing; final results announced July 2…
The developmentFirmulate’s live management experiment tested AI models’ ability to handle real company crises, revealing strengths in diagnosis but weaknesses in execution and trust.

Implications for AI in Business Management

This experiment demonstrates that current AI models excel at identifying problems but often fall short in the critical areas of decision execution, trustworthiness, and consequence management. For organizations considering AI for operational roles, this highlights the importance of evaluating models not only on their diagnostic accuracy but also on their ability to complete tasks, maintain honesty, and handle real-world pressures. The findings suggest a need to develop new benchmarks that focus on management quality rather than just chat or coding performance, potentially reshaping how AI readiness is assessed in enterprise contexts.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background

Traditional AI benchmarks mainly assess technical output, such as coding accuracy or conversational quality, often neglecting how models perform in dynamic, high-stakes situations. The Firmulate experiment fills this gap by simulating a real company environment with ongoing crises, real money mechanics, and versioned decisions. Prior to this, AI evaluation relied heavily on static tests or isolated tasks, which do not reflect the complexities of managing a business under pressure. The experiment builds on recent recognition that AI systems need to handle not just information retrieval but also consequence-aware management.

Since its inception, the concept of benchmarking AI in operational settings has gained traction, but no standardized tests have yet integrated the full scope of management responsibilities. This experiment marks a significant step toward that goal, emphasizing that the true measure of AI readiness for enterprise use involves its capacity to manage trust, escalate appropriately, and complete decisions under pressure.

“The real test for AI in management isn’t just diagnosing problems; it’s completing the job without compromising trust or execution.”

— Thorsten Meyer, lead researcher

Amazon

AI crisis management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Challenges in AI Management Evaluation

While the experiment provides valuable insights, several questions remain. It is not yet clear how well these findings generalize to larger organizations or different industries. The long-term reliability of models in maintaining trust and completing complex tasks under varied conditions is still uncertain. Additionally, the impact of integrating such benchmarks into enterprise procurement processes remains to be seen, as does the development of standardized management-focused evaluation metrics.

Amazon

enterprise AI trust and execution solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Benchmark Development and Adoption

Researchers and industry leaders are expected to build on these findings by developing more comprehensive benchmarks that include management and consequence-based tasks. Companies considering AI for operational roles should start conducting internal wargames and scenario testing to assess models’ ability to handle real-world pressures. Regulatory and standards bodies may also begin proposing new evaluation criteria focused on trustworthiness, decision completeness, and escalation protocols. The ongoing refinement of these benchmarks will determine how quickly and effectively AI can be integrated into core management functions.

Amazon

AI performance evaluation benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of the Firmulate experiment?

The experiment aims to evaluate AI models’ ability to manage a simulated company crisis, focusing on decision-making, trust, and execution, rather than just diagnostic accuracy.

Why are current benchmarks insufficient for management tasks?

Existing benchmarks primarily measure technical performance like coding or chat quality, neglecting critical management skills such as completing decisions, maintaining trust, and managing consequences under pressure.

What does the experiment reveal about AI safety and trust?

The models effectively resist manipulation attempts, such as fake messages, but often fail to follow through on decisions or escalate issues properly, exposing gaps in trust and execution.

How might this impact AI deployment in businesses?

Businesses may need to adopt new evaluation standards that test AI models’ management capabilities, ensuring they can handle real-world pressures without compromising trust or effectiveness.

What are the next steps for AI benchmarking?

Developing more comprehensive, management-focused benchmarks and conducting internal scenario tests will be key to advancing AI readiness for operational roles.

Source: ThorstenMeyerAI.com

You May Also Like

Turn Website Visitors Into Qualified Leads With AI Contact Widgets

New AI-powered contact widget enables B2B SaaS sites to automatically qualify leads by asking intent, budget, and timeline, saving sales teams time.

Appointment no-show recovery planner for therapy practices

A new appointment no-show recovery planner for small therapy practices is being tested to reduce missed appointments and improve scheduling efficiency.

Optimizing Agency Revenue With Flexible Billing Strategies

Agencies are experimenting with blended retainer-plus-usage billing to streamline invoicing and reduce revenue leakage, according to IdeaNavigator AI.

Empowering Gender Transition Through Cutting-Edge Voice Biofeedback

A mobile app using real-time acoustic biofeedback aims to support transgender individuals in voice feminization and masculinization, filling a key care gap.