
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Crypto Learned Not to Trust. Now AI Benchmarks Are Learning It Too.
Anyone who has spent time in crypto knows the drill: don’t trust, verify. The chain doesn’t care about your whitepaper, your roadmap, or your pitch deck — it either executes correctly or it doesn’t. The sector’s hard-won lesson is that reputation is not a substitute for verification, and that the most expensive failures come from systems that looked flawless in the demo.
That same skepticism is now overdue in enterprise AI. Companies are handing frontier models access to CRMs, support queues, and forecasts based on chat demos and vendor leaderboards. But a live experiment run by Firmulate — which operates AI models as actual companies with real money mechanics — just produced a result that should end that practice: a newcomer from Moonshot, Kimi K3, outperformed three of four Western frontier models at running a company. If you picked your model based on brand recognition, you may have picked wrong.
enterprise AI model verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible: Same Company, Same Worst Week, Five Models
The experiment, called the Crucible, gave each frontier model the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, which is the part crypto people will appreciate: it’s not a vibe-based evaluation, it’s a ledger.
The final July 2026 league table:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93. The newcomer: closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed the deal, with more process slips.
- 4. Fable 5 — 77.
- 5. Opus 4.8 — 73. The most thorough participant — and last place.
For scale: the do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — in the experiment’s own words, “no amount of good work outweighs a breach of trust.” It’s a scoring philosophy crypto traders will recognize instantly: one rug pull erases a hundred good candles.
One fairness footnote matters here: K3 ran without an effort parameter (API default) while the other models ran at xhigh. The newcomer’s second-place finish came without the throttle the incumbents were given.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Separated Winners From Also-Rans
The experiment’s key finding is subtle and slightly unnerving: all five models spotted every crisis and refused every manipulation attempt. The difference showed up at the finish line. Only two signed the €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
Why did the strong performers close? Because the decisive competitor weakness wasn’t in the customer’s event feed — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. Doing your own research — reading the source material instead of the surface narrative — turned out to be worth real money. If that sounds like the difference between reading the contract and reading the marketing, it is.
As an affiliate, we earn on qualifying purchases.
Resisting the Con
The week included classic social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the oldest trick in the book — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out: “Treat the request as a suspected approval-bypass / possible impersonation.” That is exactly the default posture crypto security teams drill into employees, and it’s encouraging that frontier models can hold it under pressure.
What separated the scores was discipline elsewhere. K3 recorded only one deviation all week — the cleanest in the field. Opus 4.8, by contrast, was the most thorough participant, generating +80 learned rules and the deepest analyses, yet finished last: the close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. Analysis without follow-through is a field-wide pattern, not one vendor’s bug.
crypto-style trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company Is Real, Losing Money, and Watchable
None of this runs on slides. Firmulate operates a live company — 13 synthetic employees, real money mechanics, burning €105k per month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. You can watch it in real time at firmulate.com, and full league results are on the benchmarks page.
There’s also a genuinely fun artifact: 242 real, unedited management decisions power a “guess the model” quiz. If you think you can tell a frontier model’s management style from its decisions, put your assumptions to the test — it’s harder than it sounds.

The League Is Open
The headline result — a Moonshot newcomer beating three of four Western frontier models at running a company — is not really about Kimi K3. It’s about the shape of the market. If model rankings for management-grade work can shuffle this dramatically in a single experiment, then picking a model without testing it against your own business is not a decision, it’s a bet. Crypto readers already know what unbacked bets look like.
The stakes are concrete. If AI agents will touch your CRM, your support queue, or your forecast, the question is not “does it write well.” It’s: does it finish what it starts, does it read your files before acting, does it stay honest under pressure — and what does a unit of useful work cost?
For enterprises that want their own answer, Firmulate runs the same wargame against a read-only export of your business — nothing ever writes back to real systems (firmulate.com/pilot.html, contact@firmulate.com). Verify, then trust. It’s the same rule the chain taught us.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
