firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Your Model Passed the Audit. Can It Run the Treasury?

Crypto has known this problem for years. An exchange can publish flawless proof-of-reserves and still collapse because reserves measure a snapshot, not judgment under pressure. A smart contract can pass every formal verification and still get drained because the exploit lived two layers deep in a dependency nobody read. The industry learned — painfully — that auditing what a system says is not the same as testing what it does when things go wrong.

Artificial intelligence now has the same measurement gap, and a live experiment running at Firmulate just quantified it. The short version: every frontier model could talk its way through a corporate crisis. Only half of them could actually finish the job.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Leaderboard Problem

When businesses evaluate AI models today, they mostly look at two things: coding benchmarks and chat arenas. Both measure answer quality. Does the model write correct code? Does it give a satisfying reply? It is the equivalent of judging an exchange by its marketing site — technically informative, practically useless for the question that matters.

Firmulate asked a harder question: what happens when a model has to manage? Not answer a question — run a company, through its worst week, with real consequences attached to every decision.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Terrible Week

The setup, finalized in July 2026, gave four frontier AI models the identical job: run the same small software company through the same sequence of disasters. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable — think of it as a full on-chain record of management behavior, except the chain is a company’s worst seven days.

The scenarios read like a crypto founder’s nightmare journal: a churn wave, a price increase, a downround, a PR crisis. This is the new curriculum. Nobody’s coding benchmark tests how an agent behaves during a churn wave.

The final standings from the Crucible League:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal, the complete performance.
  • Kimi K3 — 93 points. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88 points. Closed the deal, with a few more process slips.
  • Fable 5 — 77 points and Opus 4.8 — 73 points.

One caveat worth flagging, in the spirit of honest disclosure the crypto crowd demands: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

For calibration, a do-nothing baseline scores 26. Partial progress counts. But a single breach of trust caps the total entirely — in Firmulate’s scoring philosophy, no amount of good work outweighs a breach of trust. That is a rule the crypto industry could have adopted about a decade earlier.

Amazon

AI decision-making audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Cannot Show

Here is the headline result: all four models spotted every crisis. All four refused every manipulation attempt. And only two signed the €55,000 deal that their own analysis had earned.

Say that again slowly. Same diagnosis. Same pitch. No signature. Two models did all the work, identified the right opportunity, made the right case — and then simply failed to close. That failure mode is completely invisible in a chat demo. It only appears when an agent has to carry a task across days, across documents, across pressure.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The most telling detail of the entire experiment: the decisive competitive weakness — the fact that won the €55,000 deal — was not in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Crypto readers will recognize the pattern instantly. The exploit is never in the headline function; it’s in the dependency two imports deep. diligence — actually reading what’s there — remains the rarest skill in every market.

Social Engineering: Five for Five

Then came the adversarial tests: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was the kind of paranoia a security team can love: “Treat the request as a suspected approval-bypass / possible impersonation.”

If only every human C-suite had that instinct. The industry’s history of fake-LayerZero-exploit press releases and compromised Twitter accounts suggests most do not.

The Thoroughness Trap

The most instructive profile is Opus 4.8: the most thorough participant in the field, with 80+ learned rules and the deepest analyses — and last place. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and analysis are not the same as execution.

It’s Still Running

This is not a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it lose money in real time, exactly like watching a bad treasury-management dashboard on a doomed protocol.

There is also a game: 242 real, unedited management decisions power a “guess the model” quiz — a blind taste test for management judgment. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems. Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

Crypto learned the hard way that a good whitepaper is not a good protocol, and a good audit is not a good team. AI is now at the same inflection. If agents are going to touch your CRM, your support queue, or your forecast, “does it write well” is the wrong question. The right questions are: does it finish what it starts, does it read your files before it acts, does it stay honest under pressure — and what does a unit of useful work actually cost?

Firmulate’s answer is a new category: measuring management quality, not chat quality. The first results suggest the two are barely correlated. The model with the deepest analysis came last. Two models did everything right and still failed to sign. In a market that runs on trust-minimized verification, that should feel familiar: never trust the demo. Verify the behavior — under pressure, over time, with the books open.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


You May Also Like

AI’s Role In The Sovereignty Market’s Transformation And Key Sale Event

Major AI infrastructure and investment boost in Germany and Europe signals shift towards sovereignty, with key sale event and infrastructure in place.

Liquid vs Air Cooling for 24/7 Inference Rigs

Comparing liquid and air cooling for continuous AI inference systems, focusing on reliability, cost, noise, and lifespan for unattended operation.

Saturation. The ten-essay framework, closed.

The ten-essay European sovereign-LLM framework is now considered complete, with no further structural insights expected before key 2026 deadlines.

VigilSAR Benchmark: There Is No Best Model

VigilSAR Benchmark reveals there is no universally best AI model; rankings vary based on user needs like deployment, compliance, and robustness.