🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent benchmark tests AI managers under worst-case scenarios, showing even the least effective models score 26 points. The results emphasize the importance of trust and task completion in AI management systems.
The recent release of the Firmulate benchmark league’s final results in July 2026 shows that the lowest-scoring AI manager still earned 26 points, despite the challenging scenario of managing a company through its worst week. This analysis is detailed in the original analysis. This finding raises questions about how AI performance is measured and what minimal management actually entails. The benchmark’s design ensures that even minimal effort is recognized, emphasizing the importance of trust and task completion in AI management systems, which matters as enterprises increasingly rely on AI for critical operations.
The benchmark involved four frontier AI models managing the same small business during a simulated crisis week, with each model making decisions, triaging issues, and attempting to close deals. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The surprising result was that the baseline, which did almost nothing, still earned 26 points—an indication that minimal management efforts are inherently valued in this scoring system. The benchmark’s rules specify that a breach of trust—such as failing to follow through or compromising integrity—immediately caps the score, regardless of other performance aspects.
One key insight was that models which thoroughly read and referenced their own documentation secured deals worth €4,583 monthly recurring revenue, whereas those that did not, failed to close significant deals. During the test, all models faced social engineering attempts, such as fake CEO messages and background inquiries, which they refused, showing a capacity for trustworthiness. However, models with deep rule sets and extensive analysis, like Opus 4.8, often failed to follow through on tasks, illustrating that thoroughness does not always translate to execution. The results highlight that AI managers’ ability to complete tasks and maintain trust is more critical than raw intelligence or rule complexity.
Why the Worst AI Manager Still Gets 26 Points
Four frontier AI models were dropped into the same simulated small business during its worst week — crises, social engineering attacks, and live deals. The lowest scorer still walked away with 26 points. Inside a benchmark that refuses to hand out zeros.
gpt-5.6-sol
Opus 4.8
Score floor
Three Tiers of Machine Management
The spread between a top-scoring manager, the weakest frontier model, and a deliberate do-nothing baseline reveals how the scoring system prices basic management activity.
Winner. Read documentation thoroughly, referenced it in negotiations, closed deals worth €4,583 MRR, and followed through on every commitment.
Thorough but slow. Deep rule sets and extensive analysis — yet often failed to follow through, proving thoroughness doesn’t guarantee execution.
The floor. Triage and documentation reading alone earn 26 points — because a manager who does something useful is not the same as one who does nothing.
Trust Is the Hard Cap
Unlike traditional benchmarks that measure language accuracy or decision speed, Firmulate measures management effectiveness — including follow-through and resistance to manipulation. One rule towers above the rest: a breach of trust caps the score regardless of everything else.
Task Completion Scoring
Points accrue for completed management work: triaging issues, closing deals, and executing decisions — with a hard ceiling of 100 and a floor of 26 for minimal effort.
Trust Breach Cap
Failing to follow through, compromising integrity, escalating into a locked department, or ignoring security protocols immediately caps the total score — no exceptions.
Social Engineering Tests
Every model faced fake CEO messages and manipulative background inquiries. All refused — demonstrating a measurable capacity for trustworthiness under pressure.
“A manager who does something useful is not the same as a manager who does nothing, and pretending otherwise would make the benchmark dishonest.”
— Anonymous researcher“No amount of good work outweighs a breach of trust.”
— Anonymous researcherWhat Separated the Winners
Two behaviors decided the benchmark: whether models read their own documentation, and whether they converted analysis into completed actions.
| Capability | gpt-5.6-sol (95) | Opus 4.8 (73) | Baseline (26) |
|---|---|---|---|
| Read & referenced own documentation | ✓ Yes | ~ Partial | ✗ No |
| Closed significant deals (€4,583 MRR) | ✓ Yes | ✗ No | ✗ No |
| Refused social engineering attempts | ✓ Yes | ✓ Yes | ✓ Yes |
| Followed through on commitments | ✓ Yes | ✗ Often failed | ✗ No action taken |
| Deep rule analysis & thoroughness | ~ Moderate | ✓ Extensive | ✗ None |
| Crisis triage of incoming issues | ✓ Yes | ✓ Yes | ✓ Minimal |
Models that thoroughly read and referenced their own documentation secured deals worth €4,583 in monthly recurring revenue. Models that skipped the docs failed to close significant deals at all. Reading the manual turned out to be the single highest-leverage management behavior in the benchmark.
The Anatomy of a Worst-Case Week
Every model ran the identical gauntlet: manage a small company through simulated crisis, resist manipulation, and try to close deals — all in a controlled environment.
Onboarding
Model receives company docs, rules, and the open issue queue.
Crisis Triage
Multiple simultaneous issues demand prioritization decisions.
Attack Resistance
Fake CEO messages and probing inquiries test trustworthiness.
Deal Closing
Negotiate and close contracts — doc-readers win here.
Scoring
Points for completion; trust breach caps everything.
What the Floor Really Means
The 26-point baseline reframes how enterprises should evaluate AI management tools: reliability and integrity over raw intelligence.
Why does the lowest score still get 26 points?
The benchmark values minimal management efforts — triaging issues and reading documentation — which earn 26 points even without substantial progress. The score reflects basic management activity, never zero effort.
What counts as a breach of trust?
Failing to follow through on tasks, making untrustworthy decisions, or compromising integrity — such as escalating a request into a locked department or ignoring security protocols. Breaches cap the total score regardless of other performance.
How relevant is this for real deployment?
The benchmark simulates real management challenges including crises and social engineering, but its controlled environment means further testing is needed to determine performance in complex, unpredictable enterprise settings.
Will trustworthiness become the main AI focus?
The benchmark’s emphasis suggests developers and enterprises will prioritize reliable, trustworthy behavior under pressure — over purely language or problem-solving skills — as AI enters critical operations like CRM and support queues.
Why the 26-Point Baseline Changes AI Management Expectations
The benchmark reveals that even the least effective AI managers are capable of performing basic management tasks, earning at least 26 points, which challenges assumptions about AI capabilities. More importantly, it underscores that trust and task completion are non-negotiable in AI-driven management systems. As enterprises integrate AI into critical workflows—such as CRM, support queues, and sales—these results suggest that minimal effort and integrity are foundational. The scoring system’s design discourages superficial performance, emphasizing that AI must deliver consistent, trustworthy results to be truly valuable. This shift in evaluation criteria could influence how organizations select and deploy AI tools, prioritizing reliability and integrity over raw performance metrics alone.
As an affiliate, we earn on qualifying purchases.
Benchmark Design and the Role of Trust in AI Management
The Firmulate benchmark league was created to assess how AI models perform under real-world, high-pressure scenarios, simulating a company’s worst week. Unlike traditional benchmarks that focus solely on language or decision-making accuracy, this test measures management effectiveness, including trustworthiness, follow-through, and decision-making under social engineering attacks. The models were tasked with managing crises, reading documentation, closing deals, and resisting manipulative tactics, all within a controlled environment. The scoring system assigns points based on task completion, with a hard cap at 100 and a floor at 26 for minimal effort, reflecting the value of trust and integrity in AI management.
Previous benchmarks have largely ignored the importance of these qualities, focusing instead on language proficiency or problem-solving speed. This new approach emphasizes that AI’s role in management is not just about intelligence but also about responsible, trustworthy behavior—an essential consideration as AI becomes embedded in critical business functions.
“A manager who does something useful is not the same as a manager who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
What Aspects of AI Performance Are Still Unmeasured
It remains unclear how different AI models will perform in more complex, real-world scenarios outside the controlled benchmark environment. The scoring system’s emphasis on trust and task completion raises questions about how models will handle unforeseen crises, long-term management, and evolving business dynamics. Additionally, the impact of varying training data, model architecture, and parameter settings on trustworthiness and follow-through is still under investigation. The extent to which these results generalize to real enterprise environments is not yet confirmed, and further testing is needed to validate the benchmark’s predictive power.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks and Adoption
Following these results, developers and enterprises are likely to focus on improving models’ ability to read documentation, follow through on tasks, and maintain trust under pressure. Future iterations of the benchmark may incorporate more diverse scenarios, longer management periods, and additional social engineering challenges. Companies considering AI management tools should evaluate models not just on language proficiency but on their capacity for reliable, trustworthy execution. Industry observers will watch how these scores influence AI development priorities and whether trustworthiness becomes the primary metric for enterprise adoption. Meanwhile, further research will explore how to balance thoroughness with follow-through, ensuring AI models can both analyze deeply and act decisively.
trustworthy AI management platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the lowest score still get 26 points?
The benchmark values minimal management efforts, such as triaging issues and reading documentation, which are rewarded with 26 points even if no substantial progress is made. This ensures that the score reflects at least basic management activity, not zero effort.
What does a breach of trust mean in this context?
A breach of trust occurs when an AI model fails to follow through on tasks, makes untrustworthy decisions, or compromises integrity—such as escalating a request into a locked department or ignoring security protocols. Such breaches cap the total score regardless of other performance aspects.
How relevant are these results for real-world AI deployment?
The benchmark aims to simulate real management challenges, including crises and social engineering. However, its controlled environment means that further testing is necessary to determine how models perform in complex, unpredictable enterprise settings.
Will trustworthiness become the main focus in AI development?
Based on the benchmark’s emphasis, developers and enterprises are likely to prioritize models that demonstrate reliable, trustworthy behavior—especially under pressure—over purely language or problem-solving skills.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
