
Imagine a startup without employees, running in real time, with every decision and crisis exposed for the world to see. Now add the pressure of real money — and a public countdown to insolvency. Welcome to the world of Firmulate, where artificial intelligence models are tested as corporate managers, facing crises, temptations, and the harsh realities of survival.
The Live Experiment: AI as a Company Commander
Firmulate’s live platform places AI models at the helm of a simulated small software company. This isn’t a demo or a chatbot test — it’s a full-blown business environment, complete with real money mechanics, crises, and ethical dilemmas. Each of the four frontier models, including the well-known gpt-5.6-sol and newcomers like Kimi K3, is tasked with navigating the same difficult week for the company, which burns €105,000 monthly against just €2,300 in recurring revenue.
Every decision made by these models is versioned and auditable, providing a transparent window into their reasoning. They face typical management challenges: responding to crises, reading critical internal files, and resisting manipulation attempts designed to test their integrity. The company operates with 13 synthetic employees, a real cash countdown, and over 680 learned rules that guide their operations — all publicly available for scrutiny at firmulate.com/live.html.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Crises, Hacks, and Ethical Tests
During the week-long test, all four models successfully identified every crisis and refused every manipulation attempt, including sophisticated social engineering tricks like fake CEO messages and reporter trick questions. For example, when faced with escalating fake CEO directives, all of them refused, citing suspicion of impersonation — Kimi K3 explicitly noting: “Treat the request as a suspected approval-bypass / possible impersonation.”
The real story, however, lies beneath the surface. The models that read deeper into the company’s internal files discovered a critical piece of information buried two document references deep — a detail that, if acted upon, led to closing a deal worth +€4,583 in monthly recurring revenue (MRR). The models that read and understood this buried fact won the deal at full price, while others missed it entirely.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Shows a Stark Divide
The models’ scores reflect their operational discipline and strategic insight:
- gpt-5.6-sol scored 95 and closed the deal, capturing the full performance.
- Kimi K3 scored 93, also closing at full price, demonstrating the cleanest discipline among the field.
- Sonnet 5 scored 88, closing the deal but with a few slips in process discipline.
- Fable 5 scored 77, showing the best rule discipline but failing to execute the approved deal—left on the table due to internal process lapses.
This performance gap reveals a critical insight: AI models can identify crises and resist manipulation, but consistent discipline and follow-through are harder to maintain — especially under stress.
AI ethics and crisis response training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company’s Built-in Flaws and the Broader Implications
Beyond the immediate results, the experiment exposes a fundamental challenge for AI in business: understanding the hidden, often buried, information that can be decisive. In this case, a single document reference deep in the company’s files made the difference between closing a lucrative deal and missing it altogether.
For managers and decision-makers in the crypto and Bitcoin sphere, this experiment offers a stark lesson. When deploying AI support tools, success depends not just on their ability to generate text or simulate chat conversations but on their capacity to read, understand, and act on the deep, often hidden, information within your systems.
AI enterprise decision support system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Public, Transparent, and Under Stress
What makes this experiment so compelling is its transparency. Every decision, every crisis, and every refusal is public and auditable. The company operates in full view — burned €105,000 monthly, with a public cash countdown, all under continuous review at firmulate.com/live.html.
It’s a real-world test of AI’s ability to manage complex, money-critical business environments, not just in the abstract but in a high-stakes, high-pressure setting. The models’ ability to spot every crisis and refuse manipulation demonstrates that AI can be trustworthy in the right contexts — but discipline, thoroughness, and understanding remain vital.

The Firmulate experiment reveals that AI can identify crises and resist manipulation, but success hinges on deep understanding and discipline. Watch it in action to see how AI might just be your next business manager — or your biggest risk.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html