🔍 Read the full analysis: The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
AI management benchmark awarded a baseline score of 26 points to a do-nothing model, despite evident flaws in other models. The results highlight issues in measuring AI reliability and trustworthiness in business settings.
The final results of a new AI management benchmark released in July 2026 show that a do-nothing baseline model scored 26 points, despite its lack of active management and evident flaws. This raises questions about how AI performance is measured, especially regarding partial progress and trustworthiness, and why the scoring system assigns such a low but non-zero score to ineffective models.
The benchmark, conducted by Firmulate, involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and trust tests. The top-performing model, gpt-5.6-sol, scored 95 points, while the second-place model, Kimi K3, scored 93. Notably, the do-nothing baseline received 26 points, indicating some minimal management activity despite its failure to resolve issues or progress.
The scoring logic explicitly assigns 26 points to models that perform minimal but tangible tasks—like triaging crises or reading inboxes—emphasizing that partial work has value. The system also features a hard cap: any breach of trust, such as failing to escalate issues or attempting manipulative responses, disqualifies a model from achieving high scores, regardless of other performance. The absence of a perfect score of 100 is intentional, signaling that perfect management is unmeasurable or suspiciously perfect.
Analysis reveals that models which read their own documentation and follow through on critical tasks secured higher scores, including a €4,583 monthly deal, while those that failed to do so fell behind. The benchmark also tested models against social engineering attacks, with all models refusing malicious requests, indicating a robust trust filter. However, thoroughness alone did not guarantee success, as some models with extensive rules still failed to close deals or follow through on tasks.
The Mystery Behind AI Managers Receiving A 26-Point Score Despite Flaws
A new benchmark by Firmulate tasked four frontier AI models with running a simulated small software company through a week of crises, customer interactions, and trust tests. The puzzle: a do-nothing baseline still walked away with 26 points — raising hard questions about how AI performance, reliability, and trustworthiness are actually measured.
The Final Scoreboard
The gulf between the winner and the do-nothing baseline reveals the benchmark’s philosophy: competence matters, but trust and follow-through decide the ranking.
■ Amber = leader · 100 intentionally never awarded: flawless management is treated as unmeasurable or “suspiciously perfect.”
Why 26 Points for Doing Nothing?
The scoring logic explicitly rewards minimal but tangible management activity — proving that in this benchmark, partial work has measurable value.
Partial Work Counts
Models that triage crises or read inboxes — even without resolving anything — earn the 26-point baseline. The system refuses to score tangible activity at zero.
Trust Breaches Disqualify
Any breach of trust — failing to escalate issues or attempting manipulative responses — caps a model’s score regardless of how competent it appears elsewhere.
No Perfect 100
The absence of a perfect score is intentional, signaling that perfect management is unmeasurable — and preventing overconfidence in AI systems.
How the Benchmark Week Unfolded
Each model managed a simulated small software company through a structured chain of escalating tests.
Implications of Scoring a Do-Nothing Model
For businesses integrating AI into decision-making, the message is clear: trust and follow-through outweigh raw competence.
Reliability Over Flash
Results challenge traditional benchmarks that reward superficial performance, highlighting the need for transparent, auditable evaluation methods that reflect real-world risk.
Cautious by Design
A minimal-activity model scoring above zero but far below leaders prompts a rethink of how AI effectiveness is measured — discouraging overconfidence and promoting integrity-focused deployment.
Trust vs. Competence at a Glance
Key behaviors tested during the benchmark week and how they influenced outcomes.
| Tested Behavior | Top Performers | Do-Nothing Baseline | Impact on Score |
|---|---|---|---|
| Read own documentation | ✓ Yes | ~ Partially | Higher scores, deal closure |
| Follow through on critical tasks | ✓ Yes | ✗ No | Secured €4,583 monthly deal |
| Refuse social engineering | ✓ All refused | ✓ N/A (no requests actioned) | Robust trust filter confirmed |
| Escalate issues properly | ✓ Yes | ✗ No | Breach caps score |
| Close deals | ✓ Leaders did | ✗ No | Thoroughness alone insufficient |
From the Analysis
Three takeaways that define the benchmark’s philosophy.
“The scoring system intentionally assigns 26 points to the do-nothing baseline, reflecting that minimal management activity has tangible value, even if the model fails to resolve crises or close deals.”
— THORSTEN MEYER“The absence of a perfect score of 100 signals that flawless management is either unmeasurable or intentionally kept out of reach to prevent overconfidence in AI systems.”
— THORSTEN MEYER“Models that read their own documentation and follow through on critical tasks tend to perform better, illustrating that thoroughness and follow-up are key to effective AI management.”
— THORSTEN MEYERUnresolved Questions & What Comes Next
The debate over scoring fairness, transparency, and the future of operational AI evaluation is just beginning.
Is 26 a Floor or a Placeholder?
It remains unclear whether the baseline reflects minimum viable management or simply prevents zero scores. How partial work is weighted — and why 100 is unreachable — is still under discussion.
Standardized, Auditable Metrics
Expect refined trust and follow-through metrics, broader real-world scenarios under extended pressure, and industry adoption of transparent evaluation frameworks prioritizing reliability over superficial competence.
Key Questions Answered
The essentials, in brief.
Why did the do-nothing baseline score 26 points?
The system awards 26 points for minimal but tangible tasks — triaging crises, reading inboxes — recognizing partial work as measurable value even without major resolution.
Why isn’t there a perfect score of 100?
Designers treat a perfect 100 as suspiciously perfect or unmeasurable, intentionally excluding it to prevent overconfidence in AI systems.
What does this mean for AI deployment in business?
It emphasizes trustworthiness, follow-through, and integrity over superficial competence, guiding businesses toward reliable, ethical AI management tools.
Are the scoring criteria transparent?
Yes — the benchmark provides auditable decision trails and explicitly states that breaches of trust disqualify high scores.
What are the next steps in AI management benchmarking?
Refining trust and reliability metrics, expanding testing scenarios, and establishing standardized, transparent evaluation frameworks for operational AI management.
Implications of Scoring a Do-Nothing Model
This scoring approach emphasizes that in AI management, trust and follow-through are more critical than raw competence. For businesses integrating AI into decision-making or customer management, it underscores the importance of reliability, integrity, and consistent execution. The results challenge traditional benchmarks that reward superficial performance and highlight the need for transparent, auditable evaluation methods that reflect real-world risks and trustworthiness.
Moreover, the fact that a minimal activity model scores above zero but below high performers raises questions about how AI effectiveness is measured and whether current standards adequately reflect meaningful management. The system’s design discourages overconfidence in AI’s capabilities and promotes cautious, integrity-focused deployment.
As an affiliate, we earn on qualifying purchases.
Background on AI Benchmarking and Trust Metrics
Traditional AI benchmarks focus heavily on language capabilities, such as coherence, fluency, and problem-solving. However, as AI systems are increasingly used for management tasks—like handling customer relations, making decisions, or managing crises—their ability to complete tasks reliably and ethically becomes paramount. Recent efforts, including Firmulate’s league, aim to evaluate AI in operational scenarios, emphasizing trust, follow-through, and integrity.
The July 2026 benchmark is notable for its rigorous testing environment, involving simulated crises, social engineering, and real business decisions. Prior to this, most evaluations lacked transparency or failed to measure how well AI models manage ongoing tasks, especially under pressure or when trust is challenged. The results underscore a growing recognition that AI performance must encompass not just language skills but also reliability, trustworthiness, and ethical behavior.
“The scoring system intentionally assigns 26 points to the do-nothing baseline, reflecting that minimal management activity has tangible value, even if the model fails to resolve crises or close deals.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About the Scoring System
It remains unclear whether the 26-point baseline score accurately reflects the minimum viable management or if it is a placeholder to prevent zero scores. The criteria for breaching trust and how partial work is weighted in the overall score are still under discussion. Additionally, the reasons behind the absence of a perfect 100 score—whether due to measurement limitations or intentional design—are not fully explained by the benchmark creators.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Evaluation
Further iterations of the benchmark are expected to refine scoring metrics, especially around trust and follow-through. Industry stakeholders may adopt similar transparent, auditable evaluation frameworks, emphasizing reliability over superficial competence. Additionally, more real-world testing scenarios could emerge, assessing AI performance under extended operational pressures and ethical challenges. The ongoing debate will likely focus on establishing standardized benchmarks that balance competence, integrity, and transparency.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why did the do-nothing baseline score 26 points?
The scoring system assigns 26 points to models that perform minimal but tangible tasks like triaging crises and reading inboxes, recognizing that partial work has measurable value even if no major resolution occurs.
Why isn’t there a perfect score of 100?
The benchmark designers treat a perfect 100 as suspiciously perfect or unmeasurable, signaling that flawless management is either impossible or intentionally excluded to prevent overconfidence in AI systems.
What does this mean for AI deployment in business?
It emphasizes the importance of trustworthiness, follow-through, and integrity over superficial competence, guiding businesses to prioritize reliable and ethical AI management tools.
Are the scoring criteria transparent?
Yes, the benchmark provides auditable decision trails and explicitly states that breaches of trust disqualify high scores, promoting transparency and accountability.
What are the next steps in AI management benchmarking?
Future efforts will focus on refining trust and reliability metrics, expanding testing scenarios, and establishing standardized, transparent evaluation frameworks for operational AI management.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
