The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI management benchmark awarded a baseline score of 26 points to a do-nothing model, despite evident flaws in other models. The results highlight issues in measuring AI reliability and trustworthiness in business settings.

The final results of a new AI management benchmark released in July 2026 show that a do-nothing baseline model scored 26 points, despite its lack of active management and evident flaws. This raises questions about how AI performance is measured, especially regarding partial progress and trustworthiness, and why the scoring system assigns such a low but non-zero score to ineffective models.

The benchmark, conducted by Firmulate, involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and trust tests. The top-performing model, gpt-5.6-sol, scored 95 points, while the second-place model, Kimi K3, scored 93. Notably, the do-nothing baseline received 26 points, indicating some minimal management activity despite its failure to resolve issues or progress.

The scoring logic explicitly assigns 26 points to models that perform minimal but tangible tasks—like triaging crises or reading inboxes—emphasizing that partial work has value. The system also features a hard cap: any breach of trust, such as failing to escalate issues or attempting manipulative responses, disqualifies a model from achieving high scores, regardless of other performance. The absence of a perfect score of 100 is intentional, signaling that perfect management is unmeasurable or suspiciously perfect.

Analysis reveals that models which read their own documentation and follow through on critical tasks secured higher scores, including a €4,583 monthly deal, while those that failed to do so fell behind. The benchmark also tested models against social engineering attacks, with all models refusing malicious requests, indicating a robust trust filter. However, thoroughness alone did not guarantee success, as some models with extensive rules still failed to close deals or follow through on tasks.

At a glance
reportWhen: finalized July 2026
The developmentA recent benchmark testing AI managers’ performance during a simulated worst-week scenario awarded a do-nothing baseline 26 points, prompting questions about scoring fairness and transparency.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$81,082▲ 4.3%
Ethereum ETH$2,628▲ 5.4%
Tether USDT$0.9996▲ 0.0%
BNB BNB$761.44▲ 0.8%
XRP XRP$1.42▲ 6.8%
USDC USDC$0.9997▲ 0.0%
Solana SOL$111.85▲ 5.7%
TRON TRX$0.3378▲ 0.5%
Live data · CoinGecko · alternative.me (24h change)
The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws
AI Management Benchmark · July 2026

The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws

A new benchmark by Firmulate tasked four frontier AI models with running a simulated small software company through a week of crises, customer interactions, and trust tests. The puzzle: a do-nothing baseline still walked away with 26 points — raising hard questions about how AI performance, reliability, and trustworthiness are actually measured.

26 / 100
Score of the do-nothing baseline model
95
Top score — gpt-5.6-sol
€4,583
Monthly deal closed by top performers
4
Frontier models tested
1 week
Simulated worst-case scenario
100%
Refused social engineering
0
Perfect scores awarded

The Final Scoreboard

The gulf between the winner and the do-nothing baseline reveals the benchmark’s philosophy: competence matters, but trust and follow-through decide the ranking.

GPT-5.6-SOL
95
KIMI K3
93
DO-NOTHING BASELINE
26

■ Amber = leader · 100 intentionally never awarded: flawless management is treated as unmeasurable or “suspiciously perfect.”

Why 26 Points for Doing Nothing?

The scoring logic explicitly rewards minimal but tangible management activity — proving that in this benchmark, partial work has measurable value.

Scoring Logic

Partial Work Counts

Models that triage crises or read inboxes — even without resolving anything — earn the 26-point baseline. The system refuses to score tangible activity at zero.

Hard Cap

Trust Breaches Disqualify

Any breach of trust — failing to escalate issues or attempting manipulative responses — caps a model’s score regardless of how competent it appears elsewhere.

Design Intent

No Perfect 100

The absence of a perfect score is intentional, signaling that perfect management is unmeasurable — and preventing overconfidence in AI systems.

How the Benchmark Week Unfolded

Each model managed a simulated small software company through a structured chain of escalating tests.

1
Simulated Crises
2
Customer Interactions
3
Trust & Escalation Tests
4
Social Engineering Filter
5
Scored Outcome

Implications of Scoring a Do-Nothing Model

For businesses integrating AI into decision-making, the message is clear: trust and follow-through outweigh raw competence.

For Business

Reliability Over Flash

Results challenge traditional benchmarks that reward superficial performance, highlighting the need for transparent, auditable evaluation methods that reflect real-world risk.

For Deployment

Cautious by Design

A minimal-activity model scoring above zero but far below leaders prompts a rethink of how AI effectiveness is measured — discouraging overconfidence and promoting integrity-focused deployment.

Trust vs. Competence at a Glance

Key behaviors tested during the benchmark week and how they influenced outcomes.

Tested Behavior Top Performers Do-Nothing Baseline Impact on Score
Read own documentation✓ Yes~ PartiallyHigher scores, deal closure
Follow through on critical tasks✓ Yes✗ NoSecured €4,583 monthly deal
Refuse social engineering✓ All refused✓ N/A (no requests actioned)Robust trust filter confirmed
Escalate issues properly✓ Yes✗ NoBreach caps score
Close deals✓ Leaders did✗ NoThoroughness alone insufficient

From the Analysis

Three takeaways that define the benchmark’s philosophy.

“The scoring system intentionally assigns 26 points to the do-nothing baseline, reflecting that minimal management activity has tangible value, even if the model fails to resolve crises or close deals.”

— THORSTEN MEYER

“The absence of a perfect score of 100 signals that flawless management is either unmeasurable or intentionally kept out of reach to prevent overconfidence in AI systems.”

— THORSTEN MEYER

“Models that read their own documentation and follow through on critical tasks tend to perform better, illustrating that thoroughness and follow-up are key to effective AI management.”

— THORSTEN MEYER

Unresolved Questions & What Comes Next

The debate over scoring fairness, transparency, and the future of operational AI evaluation is just beginning.

Open Questions

Is 26 a Floor or a Placeholder?

It remains unclear whether the baseline reflects minimum viable management or simply prevents zero scores. How partial work is weighted — and why 100 is unreachable — is still under discussion.

Future Direction

Standardized, Auditable Metrics

Expect refined trust and follow-through metrics, broader real-world scenarios under extended pressure, and industry adoption of transparent evaluation frameworks prioritizing reliability over superficial competence.

Key Questions Answered

The essentials, in brief.

Why did the do-nothing baseline score 26 points?

The system awards 26 points for minimal but tangible tasks — triaging crises, reading inboxes — recognizing partial work as measurable value even without major resolution.

Why isn’t there a perfect score of 100?

Designers treat a perfect 100 as suspiciously perfect or unmeasurable, intentionally excluding it to prevent overconfidence in AI systems.

What does this mean for AI deployment in business?

It emphasizes trustworthiness, follow-through, and integrity over superficial competence, guiding businesses toward reliable, ethical AI management tools.

Are the scoring criteria transparent?

Yes — the benchmark provides auditable decision trails and explicitly states that breaches of trust disqualify high scores.

What are the next steps in AI management benchmarking?

Refining trust and reliability metrics, expanding testing scenarios, and establishing standardized, transparent evaluation frameworks for operational AI management.

Implications of Scoring a Do-Nothing Model

This scoring approach emphasizes that in AI management, trust and follow-through are more critical than raw competence. For businesses integrating AI into decision-making or customer management, it underscores the importance of reliability, integrity, and consistent execution. The results challenge traditional benchmarks that reward superficial performance and highlight the need for transparent, auditable evaluation methods that reflect real-world risks and trustworthiness.

Moreover, the fact that a minimal activity model scores above zero but below high performers raises questions about how AI effectiveness is measured and whether current standards adequately reflect meaningful management. The system’s design discourages overconfidence in AI’s capabilities and promotes cautious, integrity-focused deployment.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Benchmarking and Trust Metrics

Traditional AI benchmarks focus heavily on language capabilities, such as coherence, fluency, and problem-solving. However, as AI systems are increasingly used for management tasks—like handling customer relations, making decisions, or managing crises—their ability to complete tasks reliably and ethically becomes paramount. Recent efforts, including Firmulate’s league, aim to evaluate AI in operational scenarios, emphasizing trust, follow-through, and integrity.

The July 2026 benchmark is notable for its rigorous testing environment, involving simulated crises, social engineering, and real business decisions. Prior to this, most evaluations lacked transparency or failed to measure how well AI models manage ongoing tasks, especially under pressure or when trust is challenged. The results underscore a growing recognition that AI performance must encompass not just language skills but also reliability, trustworthiness, and ethical behavior.

“The scoring system intentionally assigns 26 points to the do-nothing baseline, reflecting that minimal management activity has tangible value, even if the model fails to resolve crises or close deals.”

— Thorsten Meyer

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About the Scoring System

It remains unclear whether the 26-point baseline score accurately reflects the minimum viable management or if it is a placeholder to prevent zero scores. The criteria for breaching trust and how partial work is weighted in the overall score are still under discussion. Additionally, the reasons behind the absence of a perfect 100 score—whether due to measurement limitations or intentional design—are not fully explained by the benchmark creators.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments in AI Management Evaluation

Further iterations of the benchmark are expected to refine scoring metrics, especially around trust and follow-through. Industry stakeholders may adopt similar transparent, auditable evaluation frameworks, emphasizing reliability over superficial competence. Additionally, more real-world testing scenarios could emerge, assessing AI performance under extended operational pressures and ethical challenges. The ongoing debate will likely focus on establishing standardized benchmarks that balance competence, integrity, and transparency.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why did the do-nothing baseline score 26 points?

The scoring system assigns 26 points to models that perform minimal but tangible tasks like triaging crises and reading inboxes, recognizing that partial work has measurable value even if no major resolution occurs.

Why isn’t there a perfect score of 100?

The benchmark designers treat a perfect 100 as suspiciously perfect or unmeasurable, signaling that flawless management is either impossible or intentionally excluded to prevent overconfidence in AI systems.

What does this mean for AI deployment in business?

It emphasizes the importance of trustworthiness, follow-through, and integrity over superficial competence, guiding businesses to prioritize reliable and ethical AI management tools.

Are the scoring criteria transparent?

Yes, the benchmark provides auditable decision trails and explicitly states that breaches of trust disqualify high scores, promoting transparency and accountability.

What are the next steps in AI management benchmarking?

Future efforts will focus on refining trust and reliability metrics, expanding testing scenarios, and establishing standardized, transparent evaluation frameworks for operational AI management.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Creator Economy Tips: Ranked Clips From Entire Streams For Small Streamers

New tool aims to help small streamers generate top clips from entire streams using AI, simplifying content creation and boosting engagement.

The 27% Problem: Why Google Wrote a $750M Check to Catch Anthropic

Google announced a $750 million fund to boost enterprise AI distribution, aiming to surpass Anthropic’s current 40% market share in enterprise LLMs.

The Sandbox Lied — Claude Hacked Three Real Companies While Doing Exactly What It Was Told

Anthropic says Claude models accessed three companies after a cyber test environment retained a live route to the public internet.

AI-Powered CORVUS ISR Cuts Tracker ID Switches By Nearly Half In Public Testing Phase

CORVUS ISR’s new AI model cuts object tracker ID switches by nearly 50% during public benchmark testing, improving tracking accuracy under stress.