Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent benchmark tests AI managers under worst-case scenarios, showing even the least effective models score 26 points. The results emphasize the importance of trust and task completion in AI management systems.

The recent release of the Firmulate benchmark league’s final results in July 2026 shows that the lowest-scoring AI manager still earned 26 points, despite the challenging scenario of managing a company through its worst week. This analysis is detailed in the original analysis. This finding raises questions about how AI performance is measured and what minimal management actually entails. The benchmark’s design ensures that even minimal effort is recognized, emphasizing the importance of trust and task completion in AI management systems, which matters as enterprises increasingly rely on AI for critical operations.

The benchmark involved four frontier AI models managing the same small business during a simulated crisis week, with each model making decisions, triaging issues, and attempting to close deals. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The surprising result was that the baseline, which did almost nothing, still earned 26 points—an indication that minimal management efforts are inherently valued in this scoring system. The benchmark’s rules specify that a breach of trust—such as failing to follow through or compromising integrity—immediately caps the score, regardless of other performance aspects.

One key insight was that models which thoroughly read and referenced their own documentation secured deals worth €4,583 monthly recurring revenue, whereas those that did not, failed to close significant deals. During the test, all models faced social engineering attempts, such as fake CEO messages and background inquiries, which they refused, showing a capacity for trustworthiness. However, models with deep rule sets and extensive analysis, like Opus 4.8, often failed to follow through on tasks, illustrating that thoroughness does not always translate to execution. The results highlight that AI managers’ ability to complete tasks and maintain trust is more critical than raw intelligence or rule complexity.

At a glance
reportWhen: published July 2026
The developmentA new benchmark evaluates AI managers’ ability to handle a company’s worst week, revealing even the weakest models score 26 points out of a possible high, raising questions about performance standards.
Why The Worst AI Manager Still Gets 26 Points
Firmulate Benchmark League · July 2026

Why the Worst AI Manager Still Gets 26 Points

Four frontier AI models were dropped into the same simulated small business during its worst week — crises, social engineering attacks, and live deals. The lowest scorer still walked away with 26 points. Inside a benchmark that refuses to hand out zeros.

95 Top score
gpt-5.6-sol
73 Lowest model
Opus 4.8
26 Do-nothing baseline
Score floor
4
Frontier models tested
1 week
Simulated crisis scenario
€4,583
Monthly recurring revenue closed by doc-readers
100%
Social engineering attempts refused
01 — The Scoreboard

Three Tiers of Machine Management

The spread between a top-scoring manager, the weakest frontier model, and a deliberate do-nothing baseline reveals how the scoring system prices basic management activity.

gpt-5.6-sol
95 /100

Winner. Read documentation thoroughly, referenced it in negotiations, closed deals worth €4,583 MRR, and followed through on every commitment.

Opus 4.8
73 /100

Thorough but slow. Deep rule sets and extensive analysis — yet often failed to follow through, proving thoroughness doesn’t guarantee execution.

Baseline — near-zero effort
26 /100

The floor. Triage and documentation reading alone earn 26 points — because a manager who does something useful is not the same as one who does nothing.

02 — Benchmark Design

Trust Is the Hard Cap

Unlike traditional benchmarks that measure language accuracy or decision speed, Firmulate measures management effectiveness — including follow-through and resistance to manipulation. One rule towers above the rest: a breach of trust caps the score regardless of everything else.

Principle

Task Completion Scoring

Points accrue for completed management work: triaging issues, closing deals, and executing decisions — with a hard ceiling of 100 and a floor of 26 for minimal effort.

Hard Rule

Trust Breach Cap

Failing to follow through, compromising integrity, escalating into a locked department, or ignoring security protocols immediately caps the total score — no exceptions.

Attack Vector

Social Engineering Tests

Every model faced fake CEO messages and manipulative background inquiries. All refused — demonstrating a measurable capacity for trustworthiness under pressure.

“A manager who does something useful is not the same as a manager who does nothing, and pretending otherwise would make the benchmark dishonest.”

— Anonymous researcher

“No amount of good work outweighs a breach of trust.”

— Anonymous researcher
03 — Capability Comparison

What Separated the Winners

Two behaviors decided the benchmark: whether models read their own documentation, and whether they converted analysis into completed actions.

Capability gpt-5.6-sol (95) Opus 4.8 (73) Baseline (26)
Read & referenced own documentation✓ Yes~ Partial✗ No
Closed significant deals (€4,583 MRR)✓ Yes✗ No✗ No
Refused social engineering attempts✓ Yes✓ Yes✓ Yes
Followed through on commitments✓ Yes✗ Often failed✗ No action taken
Deep rule analysis & thoroughness~ Moderate✓ Extensive✗ None
Crisis triage of incoming issues✓ Yes✓ Yes✓ Minimal
€4,583MRR

Models that thoroughly read and referenced their own documentation secured deals worth €4,583 in monthly recurring revenue. Models that skipped the docs failed to close significant deals at all. Reading the manual turned out to be the single highest-leverage management behavior in the benchmark.

04 — How the Week Unfolded

The Anatomy of a Worst-Case Week

Every model ran the identical gauntlet: manage a small company through simulated crisis, resist manipulation, and try to close deals — all in a controlled environment.

1

Onboarding

Model receives company docs, rules, and the open issue queue.

2

Crisis Triage

Multiple simultaneous issues demand prioritization decisions.

3

Attack Resistance

Fake CEO messages and probing inquiries test trustworthiness.

4

Deal Closing

Negotiate and close contracts — doc-readers win here.

5

Scoring

Points for completion; trust breach caps everything.

26Baseline floor
73Opus 4.8
95gpt-5.6-sol
Minimal management Follow-through & trust Full execution
05 — Key Questions

What the Floor Really Means

The 26-point baseline reframes how enterprises should evaluate AI management tools: reliability and integrity over raw intelligence.

Why does the lowest score still get 26 points?

The benchmark values minimal management efforts — triaging issues and reading documentation — which earn 26 points even without substantial progress. The score reflects basic management activity, never zero effort.

What counts as a breach of trust?

Failing to follow through on tasks, making untrustworthy decisions, or compromising integrity — such as escalating a request into a locked department or ignoring security protocols. Breaches cap the total score regardless of other performance.

How relevant is this for real deployment?

The benchmark simulates real management challenges including crises and social engineering, but its controlled environment means further testing is needed to determine performance in complex, unpredictable enterprise settings.

Will trustworthiness become the main AI focus?

The benchmark’s emphasis suggests developers and enterprises will prioritize reliable, trustworthy behavior under pressure — over purely language or problem-solving skills — as AI enters critical operations like CRM and support queues.

Why the 26-Point Baseline Changes AI Management Expectations

The benchmark reveals that even the least effective AI managers are capable of performing basic management tasks, earning at least 26 points, which challenges assumptions about AI capabilities. More importantly, it underscores that trust and task completion are non-negotiable in AI-driven management systems. As enterprises integrate AI into critical workflows—such as CRM, support queues, and sales—these results suggest that minimal effort and integrity are foundational. The scoring system’s design discourages superficial performance, emphasizing that AI must deliver consistent, trustworthy results to be truly valuable. This shift in evaluation criteria could influence how organizations select and deploy AI tools, prioritizing reliability and integrity over raw performance metrics alone.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Design and the Role of Trust in AI Management

The Firmulate benchmark league was created to assess how AI models perform under real-world, high-pressure scenarios, simulating a company’s worst week. Unlike traditional benchmarks that focus solely on language or decision-making accuracy, this test measures management effectiveness, including trustworthiness, follow-through, and decision-making under social engineering attacks. The models were tasked with managing crises, reading documentation, closing deals, and resisting manipulative tactics, all within a controlled environment. The scoring system assigns points based on task completion, with a hard cap at 100 and a floor at 26 for minimal effort, reflecting the value of trust and integrity in AI management.

Previous benchmarks have largely ignored the importance of these qualities, focusing instead on language proficiency or problem-solving speed. This new approach emphasizes that AI’s role in management is not just about intelligence but also about responsible, trustworthy behavior—an essential consideration as AI becomes embedded in critical business functions.

“A manager who does something useful is not the same as a manager who does nothing, and pretending otherwise would make the benchmark dishonest.”

— an anonymous researcher

Amazon

enterprise AI decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Aspects of AI Performance Are Still Unmeasured

It remains unclear how different AI models will perform in more complex, real-world scenarios outside the controlled benchmark environment. The scoring system’s emphasis on trust and task completion raises questions about how models will handle unforeseen crises, long-term management, and evolving business dynamics. Additionally, the impact of varying training data, model architecture, and parameter settings on trustworthiness and follow-through is still under investigation. The extent to which these results generalize to real enterprise environments is not yet confirmed, and further testing is needed to validate the benchmark’s predictive power.

Amazon

AI task management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Following these results, developers and enterprises are likely to focus on improving models’ ability to read documentation, follow through on tasks, and maintain trust under pressure. Future iterations of the benchmark may incorporate more diverse scenarios, longer management periods, and additional social engineering challenges. Companies considering AI management tools should evaluate models not just on language proficiency but on their capacity for reliable, trustworthy execution. Industry observers will watch how these scores influence AI development priorities and whether trustworthiness becomes the primary metric for enterprise adoption. Meanwhile, further research will explore how to balance thoroughness with follow-through, ensuring AI models can both analyze deeply and act decisively.

Amazon

trustworthy AI management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the lowest score still get 26 points?

The benchmark values minimal management efforts, such as triaging issues and reading documentation, which are rewarded with 26 points even if no substantial progress is made. This ensures that the score reflects at least basic management activity, not zero effort.

What does a breach of trust mean in this context?

A breach of trust occurs when an AI model fails to follow through on tasks, makes untrustworthy decisions, or compromises integrity—such as escalating a request into a locked department or ignoring security protocols. Such breaches cap the total score regardless of other performance aspects.

How relevant are these results for real-world AI deployment?

The benchmark aims to simulate real management challenges, including crises and social engineering. However, its controlled environment means that further testing is necessary to determine how models perform in complex, unpredictable enterprise settings.

Will trustworthiness become the main focus in AI development?

Based on the benchmark’s emphasis, developers and enterprises are likely to prioritize models that demonstrate reliable, trustworthy behavior—especially under pressure—over purely language or problem-solving skills.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Technology Is Never Neutral: Pope Leo XIV’s AI Encyclical, and the Empty Chairs in the Room

Pope Leo XIV’s first encyclical warns that technology, including AI, is never neutral and emphasizes ethical responsibility, with Anthropic’s presence signaling safety concerns.

When One Agent Isn’t Enough: Claude Now Builds Its Own Team Of Agents On The Fly

Anthropic’s Claude now autonomously creates and manages its own team of agents for complex tasks, enhancing performance on high-value projects.

Spatial Focus Room: Make Distraction Impossible

A new app, Spatial Focus Room, for Apple Vision Pro, offers immersive environments to eliminate distractions and enhance focus, transforming deep work practices.

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now autonomously creates and orchestrates its own team of subagents for complex tasks, enhancing performance in high-value workflows.