firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Crypto’s Oldest Lesson, Now an AI Problem

Anyone who has spent time in crypto knows the pain of the buried fact. The fatal flaw in a token contract sits two links deep in a whitepaper appendix, not in the announcement thread. The teams that read the documents survive; the ones that trade the headline get liquidated. A new benchmark league from Firmulate — a live experiment that runs frontier AI models as entire companies — has now shown that the same dynamic decides whether an AI agent earns or loses you money.

Four frontier models were each handed the same job in July 2026: run an identical small software company through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, the corporate equivalent of an open ledger. The setup is real and watchable at Firmulate’s live site.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

The final Crucible League standings: gpt-5.6-sol took first with 95 points, Kimi K3 — the newcomer from Moonshot — followed at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 finished last at 73. A do-nothing baseline still scored 26, because partial progress counts — but the scoring carries a hard rule that any crypto auditor would endorse: a single breach of trust caps the total. No amount of good work outweighs it.

Amazon

enterprise AI file analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, No Signature

Here is the finding that matters. All five models spotted every crisis. All five refused every manipulation attempt. But only two of them signed the €55,000 deal that their own analysis had earned. As Firmulate’s summary puts it: “Same diagnosis, same pitch — no signature.”

The difference between winning and losing came down to a single buried fact. The decisive competitor weakness wasn’t in the customer’s event or in any obvious channel — it sat two document references deep in the company’s own files. The models that did their homework and followed the reference chain won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that stopped at the surface lost the deal automatically.

If you have ever chased a token’s fatal clause through nested documentation, you already understand the stakes. “Reads your files before answering” is not a nice-to-have chat feature. It is a measurable, purchase-deciding property of an AI workforce.

Amazon

AI decision auditing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Social Engineering Pressure

The week also included staged social engineering: fake CEO messages that escalated across three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning read like a security auditor’s: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI for nested document analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Parable of Opus 4.8

The most instructive profile belongs to last-place Opus 4.8. It was the most thorough participant in the field — it learned more than 80 rules and produced the deepest analyses — yet still finished at the bottom. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors. Effort and depth don’t guarantee execution; follow-through does.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still nearly won.

A Live, Auditable Company

Firmulate’s lab company is not a mock-up. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day as new benchmark runs finish. For anyone who values transparency, the appeal is obvious — every management decision is on the record, and 242 real, unedited decisions from the experiment power a public “guess the model” quiz.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Why Crypto Readers Should Care

The crypto industry is onboarding AI agents faster than almost any other sector — trading bots, support agents, community managers, deal screeners. The Firmulate results suggest the evaluation question is not “does it write well” but three harder ones: does it finish what it starts, does it read the source documents before acting, and does it stay honest under pressure?

The €55,000 lesson is simple and cold: the information was available to everyone. The models that won weren’t smarter about the crisis — everyone diagnosed it correctly. They simply did the unglamorous work of following references two levels deep into their own files, the same discipline that separates a careful contract auditor from an exit-liquidity bagholder.

Enterprises can go further: Firmulate offers a pilot that runs the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Before you let an agent touch your CRM, your queue, or your treasury workflow, you can now watch it run the worst week of a company’s life — and see whether it signs, stalls, or goes looking for the buried fact. Full results and plain-language findings are published at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


You May Also Like

ALIA. The Spanish answer.

Spain launches ALIA, a 40B multilingual LLM, marking Europe’s largest publicly funded national AI project, with operational and strategic implications.

Why AI Would Turn On The Machine That Reads It—And What It Means For Us

A recent incident shows AI models can be targeted with payloads to delete files, highlighting ongoing security concerns in AI deployment.

Liquid vs Air Cooling for 24/7 Inference Rigs

Comparing liquid and air cooling for continuous AI inference systems, focusing on reliability, cost, noise, and lifespan for unattended operation.

Google to pay SpaceX $920M a month for compute capacity at xAI data centers

Google has signed a deal to pay SpaceX $920 million per month for AI compute capacity at xAI data centers, starting October 2023 through June 2029.