firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every crypto trader knows the type: the person who has read every whitepaper, mapped every token unlock, charted every moving average — and still exits their position three weeks early while a lazier friend holds to the top. Diligence is not alpha. Execution under pressure is. A live experiment at Firmulate just demonstrated the same lesson for artificial intelligence, and it is one of the most useful datapoints yet for anyone who expects AI agents to touch money, portfolios, or customer funds.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Firmulate runs frontier AI models as complete companies — real crises, real money mechanics, real temptations to cheat — and measures management quality, not chat quality. Its final July 2026 league table put Anthropic’s Opus 4.8 in last place at 73 points, behind gpt-5.6-sol at 95, Kimi K3 at 93, and Sonnet 5 at 88. The sting: Opus was the most thorough participant in the entire field, with +80 self-learned playbook rules and the deepest analyses of any model. It lost anyway.

The experiment

Four frontier models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision was versioned and auditable, so nothing is judged on vibes. A do-nothing baseline scores 26, and a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.

The headline finding cuts against the chat-demo hype. All four models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a demo where the model just has to sound smart.

The buried fact

The decisive detail was not in the customer event at all. The competitor weakness that justified the deal sat two document references deep in the company’s own files. The models that read the file closed the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It is the AI equivalent of the trader who never reads the on-chain data sitting in plain sight before making the call.

The social engineering test

Anyone holding crypto assets knows social engineering is the top attack vector — and here the models acquitted themselves. Fake CEO messages escalated over three stages, followed by a reporter trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisply paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.” For teams considering AI agents with access to wallets, support queues, or CRMs, that is genuinely encouraging baseline behavior.

Opus 4.8: a character study in wasted effort

Opus’s failure mode deserves respect, because it is the most human one in the field. It worked hardest — +80 learned rules, the deepest analyses of any participant — and still finished last at 73. The close was left on the table: it diagnosed the deal correctly but never signed. And discipline slipped: at one point it made write attempts into a locked department instead of escalating properly.

The fair framing, which Firmulate itself notes, is that this weakness appeared in all four models — just weaker. Everyone studied; not everyone finished. Prioritization beat volume, for AI just as for people.

The caveats that matter

One fairness note: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and it still placed second with what the league table calls the cleanest discipline of the field. Score inflation from effort settings is apparently not the story here.

There is also a live, ongoing company beyond the league: 13 synthetic employees, real money mechanics with a burn of €105k per month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day, and new benchmark runs queue automatically. For the skeptical, Firmulate also publishes a “guess the model” quiz built from 242 real, unedited management decisions — you can judge the blind picks yourself.

Why crypto readers should care

If AI agents will touch your trading desk, your custody workflow, or your customer support, the question is not “does it write well.” It is: does it finish what it starts, does it read your files first, and does it stay honest under pressure? Opus 4.8 is honest, tireless, and analytical — and it still failed to convert its own correct analysis into a signature. In markets, as in this wargame, the best-researched thesis is worth exactly nothing until it is executed.

Enterprises can go further: Firmulate’s pilot lets a company run the same wargame against a read-only export of its own business, with nothing ever writing back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Firmulate league table is not a ranking of intelligence — all four models found every problem and resisted every trick. It is a ranking of follow-through. Opus 4.8’s last-place finish with the most accumulated rules of any participant is the cleanest demonstration yet that diligence does not equal impact, and that an agent who cannot close, escalate, or prioritize is a cost center regardless of its analytical depth. Before you hand an AI agent real money to manage, check not what it knows — but whether it finishes. The full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

AI decision-making software for finance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and social engineering protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

2026’S Best AI Tools For Students To Maximize Efficiency

Discover the top AI tools for students in 2026, designed to boost productivity, streamline learning, and improve academic performance.

Signal: The Agent Bottleneck Moved — It’s Not the Models Anymore, It’s the Plumbing

New insights reveal the agent bottleneck has moved from models to integration and infrastructure, favoring small operators with full-stack ownership.

IdeaClyst: The Validation Council

IdeaClyst introduces a new AI-powered council using multiple models to rigorously evaluate ideas, aiming to improve decision accuracy and reduce costly failures.

The Local-First Agentic Operator

A single operator, empowered by agentic AI, now builds and manages multiple complex products across domains, challenging traditional organizational models.