firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Crypto people live by one rule: don’t trust, verify. A whitepaper’s claims mean nothing until every transaction is on-chain and auditable. So when an AI benchmark starts handing out suspiciously round numbers — a perfect 100, a flawless demo — the correct response is the same one you’d give a token promising 40% monthly yield: where’s the proof?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That instinct is exactly what makes the Firmulate benchmark worth a close look. In its final July 2026 “Crucible League,” four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, the way every transaction is on a public chain. And the scoring system does something most benchmarks never dare: it publishes its own floor.

A manager-bot that does nothing — no decisions, no initiative, just sits there — scores 26. Not 0. That single number tells you more about the honesty of the exercise than a hundred glossy demo videos.

Why the Floor Is 26, Not Zero

Most leaderboards are built to flatter. A model either nails a prompt or it doesn’t, and the marketing team gets a beautiful number to trumpet. Firmulate took the opposite approach: it measured management quality, not chat quality, and it scored partial progress.

Think of it like this: in crypto, a transaction that confirms is worth something even if the app around it is ugly. Useful work is useful work. A model that correctly spots a crisis, reads the files, and refuses a scam has genuinely done part of the manager’s job — even if it never closes the deal. Doing nothing but doing it correctly still avoids catastrophic downside, and the scoring acknowledges that with 26 points.

The flip side is harsher, and it’s the part every auditor will appreciate: a single breach of trust caps the total score. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” Sound familiar? It’s the same logic as a smart contract — one exploit drains the pool regardless of how many flawless blocks came before.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Note what’s absent: a round 100. The benchmark’s own designers treat a perfect score as a red flag — distrust of round 100s is baked into the philosophy. Nothing that runs a real company through a real week emerges spotless.

One fairness note, disclosed openly like a good audit trail: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Amazon

AI audit and verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Winners

Here’s the buried fact, literally. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive edge wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that only models willing to actually read the internal records would find. The models that did read it won the deal at full price, worth +€4,583 in monthly recurring revenue. In crypto terms: the alpha was in the on-chain data, not the roadmap pitch. Those who checked the source got paid.

The Social Engineering Test

The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was exactly the paranoid stance you’d want: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Trap

Opus 4.8 is the cautionary tale: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. It left the close on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four top contenders. Brilliance without follow-through scores like brilliance without follow-through. No curve, no charity.

Amazon

AI transparency benchmarking platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Company Run

Firmulate isn’t a one-off paper. There’s a live company running continuously: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. New benchmark runs queue up and the league table updates automatically. It’s watchable at firmulate.com/live, in the same spirit as a public block explorer: don’t take the results on faith, go look.

There’s also a guessing game with real stakes for your intuition: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The crypto world learned early that self-reported performance is worthless and verifiable trails are everything. AI benchmarks are now facing the same reckoning. Firmulate’s 26-point floor for doing nothing, its refusal to award a round 100, and its one-breach cap on trust are the scoring equivalent of a public audit: designed so that the numbers can embarrass their owners.

The headline finding is the one to remember. Every model talked a good game — spotted every crisis, refused every scam. Only two finished the job. That gap is invisible in chat demos and decisive in production. Before you let an AI agent near your CRM, your support queue, or your treasury, ask the question this benchmark actually answers: does it finish what it starts, and does it stay honest when finishing gets hard?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

AI trust and security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 AI Mini PCs That Will Dominate 2026

A detailed look at the 10 AI mini PCs expected to dominate in 2026, highlighting features, performance, and future-proofing for AI workloads.

AI Transparency: Anthropic’s Watermarking And Its Role In Society

Anthropic has implemented watermarking in its Claude AI system to aid content provenance. Details on the method and reliability are still emerging.

Readiness: Before You Fund the Answer

A new diagnostic tool offers companies a 20-minute assessment to determine if their AI implementation is truly ready, preventing costly failures.

The Safety Card, Played From Every Side: David Sacks, Anthropic, and the Fable Standoff

David Sacks says Anthropic ignored a serious Fable S jailbreak. Anthropic says the flaw was minor. Key evidence remains private.