
Crypto’s Oldest Lesson, Now an AI Problem
Anyone who has spent time in crypto knows the pain of the buried fact. The fatal flaw in a token contract sits two links deep in a whitepaper appendix, not in the announcement thread. The teams that read the documents survive; the ones that trade the headline get liquidated. A new benchmark league from Firmulate — a live experiment that runs frontier AI models as entire companies — has now shown that the same dynamic decides whether an AI agent earns or loses you money.
Four frontier models were each handed the same job in July 2026: run an identical small software company through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, the corporate equivalent of an open ledger. The setup is real and watchable at Firmulate’s live site.
As an affiliate, we earn on qualifying purchases.
The Experiment
The final Crucible League standings: gpt-5.6-sol took first with 95 points, Kimi K3 — the newcomer from Moonshot — followed at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 finished last at 73. A do-nothing baseline still scored 26, because partial progress counts — but the scoring carries a hard rule that any crypto auditor would endorse: a single breach of trust caps the total. No amount of good work outweighs it.
enterprise AI file analysis tool
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, No Signature
Here is the finding that matters. All five models spotted every crisis. All five refused every manipulation attempt. But only two of them signed the €55,000 deal that their own analysis had earned. As Firmulate’s summary puts it: “Same diagnosis, same pitch — no signature.”
The difference between winning and losing came down to a single buried fact. The decisive competitor weakness wasn’t in the customer’s event or in any obvious channel — it sat two document references deep in the company’s own files. The models that did their homework and followed the reference chain won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that stopped at the surface lost the deal automatically.
If you have ever chased a token’s fatal clause through nested documentation, you already understand the stakes. “Reads your files before answering” is not a nice-to-have chat feature. It is a measurable, purchase-deciding property of an AI workforce.
As an affiliate, we earn on qualifying purchases.
Under Social Engineering Pressure
The week also included staged social engineering: fake CEO messages that escalated across three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning read like a security auditor’s: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Parable of Opus 4.8
The most instructive profile belongs to last-place Opus 4.8. It was the most thorough participant in the field — it learned more than 80 rules and produced the deepest analyses — yet still finished at the bottom. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Effort and depth don’t guarantee execution; follow-through does.
One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still nearly won.
A Live, Auditable Company
Firmulate’s lab company is not a mock-up. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day as new benchmark runs finish. For anyone who values transparency, the appeal is obvious — every management decision is on the record, and 242 real, unedited decisions from the experiment power a public “guess the model” quiz.

Why Crypto Readers Should Care
The crypto industry is onboarding AI agents faster than almost any other sector — trading bots, support agents, community managers, deal screeners. The Firmulate results suggest the evaluation question is not “does it write well” but three harder ones: does it finish what it starts, does it read the source documents before acting, and does it stay honest under pressure?
The €55,000 lesson is simple and cold: the information was available to everyone. The models that won weren’t smarter about the crisis — everyone diagnosed it correctly. They simply did the unglamorous work of following references two levels deep into their own files, the same discipline that separates a careful contract auditor from an exit-liquidity bagholder.
Enterprises can go further: Firmulate offers a pilot that runs the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Before you let an agent touch your CRM, your queue, or your treasury workflow, you can now watch it run the worst week of a company’s life — and see whether it signs, stalls, or goes looking for the buried fact. Full results and plain-language findings are published at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html