firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For anyone who invests in a business, the gap between a confident forecast and a signed deal matters. Firmulate’s live experiment puts AI models in charge of the same small software company during its worst week and watches what they actually do. The result is a practical question for investors and executives alike: can an AI agent carry a decision from diagnosis through to action?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A controlled test with real stakes

In the final Crucible League, published in July 2026, five models faced the same customers, crises and temptations. Their scores ranged from 95 for gpt-5.6-sol and 93 for Kimi K3 to 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. The do-nothing baseline scored 26. The league also makes trust a hard boundary: a single breach caps the total, because “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. They reached the same diagnosis and made the same pitch; some still left without a signature. That gap between knowing what to do and doing it is the experiment’s central finding.

The clues were inside the company

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won at full price, worth +€4,583 in monthly recurring revenue. The result is a reminder that a decision can depend on whether an AI agent consults relevant company records, not just whether it responds sensibly to the latest message.

The trust tests were direct. Fake CEO messages escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its stance on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. More analysis, by itself, did not ensure a better business outcome.

The live company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. At Firmulate, readers can follow the live experiment; a quiz built from 242 real, unedited management decisions invites them to guess which model made each call.

From observing to testing your own company

For a company weighing AI agents, a league result is a starting point, not a substitute for testing against its own risks. Firmulate’s proposed enterprise pilot uses a read-only export of a company’s business to run crisis scenarios and produce a board report, including model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The aim is to let decision-makers see how models handle their own customers, records and pressure points before those systems are put to work.

One comparison deserves context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That is relevant when reading the rankings, alongside the experiment’s broader finding about execution and trust.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s live company shows what a controlled AI business wargame can reveal: sound crisis recognition does not guarantee follow-through, and important evidence may sit in internal records. Enterprises can take the next step with a pilot based on a read-only business export, with no write-back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Real-Time Business Closure Alerts: Unlock Hidden Asset Opportunities

New real-time alert system aims to help liquidation buyers identify closing businesses faster, unlocking hidden asset opportunities post-pandemic.

Creating Competitive Offers For Fractional Leaders In B2B SaaS

A new offer builder tool aims to help executives transition to fractional leadership roles by packaging services effectively, boosting first-client success.

What a Do-Nothing AI Benchmark Reveals About Trust and Performance in Business AI

Discover how a simple do-nothing AI baseline scores 26 points, setting a trustworthiness floor in business automation. Learn what honest AI behavior really costs.

AI Models Stand Firm Against Social Engineering Tests, Reaffirming Trust in Automation

AI models face social engineering tests and refuse manipulation, confirming trustworthiness and integrity—crucial for secure, reliable automation in business operations.