
A business you can watch under pressure
Crypto readers know the appeal of radical transparency: claims become more meaningful when outsiders can inspect the record. Firmulate applies a similar instinct to company building. Its software business has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue, and displays a public cash countdown as it fights for survival.
This is not a polished demonstration built around a single successful prompt. The company runs every business day, with each day’s work versioned. Its employees have accumulated more than 680 self-learned playbook rules. Visitors can watch the company live, following a continuing business story in which money, customers and operational discipline all matter.
The result is build-in-public taken to an unusually exposed conclusion. Firmulate is not merely publishing milestones after the fact. It is making the struggle itself visible: the burn, the thin revenue base, the decisions and the shrinking runway.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week, repeated under equal conditions
Firmulate also used the company as the setting for the Crucible League, completed in July 2026. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations remained the same; only the model changed. Every decision was versioned and auditable.
The final table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad result initially looks reassuring. All models identified every crisis, and all rejected every attempt at manipulation. Yet recognition was not the same as execution. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
The deal depended on reading beyond the obvious
The decisive weakness in a competitor was not presented in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding should resonate beyond model rankings. In a real organization, useful knowledge is rarely gathered into one convenient message. It sits in contracts, notes and earlier work. An AI worker may understand the immediate request while still missing the fact that changes the commercial outcome. The Firmulate result turns that distinction into a visible business consequence: the difference between preparing a credible pitch and actually securing the revenue.
Pressure tested more than salesmanship
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even under that difference, K3 finished only 2 points behind the leader and completed the deal.
Opus 4.8 produced another revealing contrast. It was the most thorough participant, adding 80 learned rules and delivering the deepest analyses, yet it finished last. It left the close on the table, and its discipline weakened when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
Thoroughness, in other words, did not guarantee completion. The experiment’s most consequential gaps appeared between knowing and doing: reading far enough, closing the deal and respecting boundaries when the straightforward route was blocked. Readers can also examine what the synthetic employees actually say, rather than relying only on the final rankings.

As an affiliate, we earn on qualifying purchases.
A public company story with real stakes
Firmulate’s live company turns AI evaluation into something closer to an unfolding operating record. The 13 synthetic employees are not judged solely on whether their language sounds persuasive. Their work meets customers, money pressure, manipulation attempts and organizational limits, while the company’s €105k monthly burn and €2.3k MRR remain visible.
For an audience accustomed to asking what can be verified, that is the central attraction. The public record does not make the company healthy, nor does it erase the distance between synthetic workers and a conventional staff. It makes performance inspectable while the consequences are still developing.
The most striking lesson from the Crucible League is therefore not that every model found the crises or rejected the traps. It is that apparently similar reasoning produced materially different business outcomes. Some models found the buried fact and secured the revenue. Others understood the situation, prepared the pitch and stopped before the signature. Firmulate makes that unfinished last mile—and the cost of it—part of the story anyone can follow.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.