AI’s Management Flaws Become Clear After Providing The Right Solution
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate’s AI experiment demonstrated that while models can understand and diagnose crises, they often fail to complete work that leads to signed deals. The findings highlight management flaws in AI deployment.

Firmulate’s live experiment has shown that AI models can correctly identify crises and formulate appropriate responses, but often fail to complete the work necessary for successful commercial outcomes. This exposes a key management flaw in deploying AI for operational decision-making, as models may understand the situation but do not reliably translate that understanding into finished, trustworthy work.

The experiment involved a simulated company with 13 synthetic employees and real financial mechanics, where AI models were tasked with diagnosing crises, resisting manipulation, and completing critical work. Despite all models correctly identifying issues and formulating responses, only two signed the €55,000 deal their analysis supported. The experiment demonstrated that the decisive factor was not understanding or reasoning, but the discipline and execution of completing the task.

Notably, the models faced social-engineering attempts, such as fake CEO messages, which all rejected. However, the most thorough model, Opus 4.8, despite extensive analysis and learning +80 rules, failed to finalize the deal when attempting to escalate into a locked department, illustrating that more analysis does not guarantee successful completion. The findings suggest that AI’s management flaws are rooted in its ability to translate understanding into action under real-world pressures.

At a glance
reportWhen: ongoing, results published July 2026
The developmentFirmulate’s live company test revealed AI models can diagnose issues but struggle to finalize work, exposing management flaws in AI decision-making.
Crypto market snapshot
Fear & Greed Index
25/100 — Extreme Fear
Bitcoin BTC$63,916▲ 1.9%
Ethereum ETH$1,843▲ 0.9%
Tether USDT$0.9992▲ 0.0%
BNB BNB$567.42▲ 0.0%
USDC USDC$0.9999▼ 0.0%
XRP XRP$1.09▲ 0.5%
Solana SOL$74.92▲ 0.6%
TRON TRX$0.3217▼ 0.1%
Live data · CoinGecko · alternative.me (24h change)

Implications of AI’s Inability to Finish Tasks Under Pressure

This experiment reveals that AI models, even when they understand the problem and formulate correct responses, can falter at the final stage of completing work that leads to revenue or operational success. For enterprises, this underscores the importance of evaluating not just AI reasoning but also its discipline and execution capabilities. The findings challenge assumptions that more thorough analysis automatically results in better outcomes, highlighting a critical gap in AI management and deployment.

Amazon

AI task management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Testing in Business Operations

Recent developments in AI have focused on improving understanding, reasoning, and safety, but less attention has been paid to how models perform when required to complete real-world tasks that impact business outcomes. Firmulate’s ongoing experiments, including live tests with simulated companies, aim to expose these management flaws. Previous benchmarks have primarily measured AI accuracy or safety, but this experiment emphasizes the importance of execution discipline in operational contexts.

“The models understood the situation consistently but failed to translate that understanding into finished, trustworthy work when it mattered most.”

— an anonymous researcher

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI’s Completion Failures

It is not yet confirmed whether these findings generalize across different industries or operational scenarios. The experiment was conducted within a simulated company environment, and real-world complexities may introduce additional challenges. Further testing is needed to determine if similar management flaws appear in live enterprise deployments under various conditions.

Amazon

AI project completion tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Evaluation and Deployment

Organizations should incorporate operational discipline tests, like those used in Firmulate’s experiments, into their AI evaluation processes. Future research may focus on developing models that better translate understanding into action, and on establishing governance frameworks to ensure AI completes critical tasks reliably. Continued live testing and benchmarking will be essential to address these management flaws.

Amazon

AI operational decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models fail to complete work despite understanding the problem?

According to recent experiments, models often lack the discipline or decision-making protocols necessary to translate understanding into action, especially under pressure or when facing real-world manipulation attempts.

What does this mean for businesses deploying AI?

It suggests that companies should evaluate not only AI’s reasoning and safety but also its ability to reliably finish tasks and close deals, which are critical for operational success.

Are these findings applicable to all AI systems?

While the experiment provides valuable insights, it was conducted in a controlled, simulated environment. Further testing in diverse real-world settings is needed to confirm the generality of these management flaws.

How can organizations improve AI’s completion performance?

Implementing operational discipline checks, designing models with clear decision and escalation protocols, and conducting live scenario testing can help improve AI’s ability to finish critical work.

Source: ThorstenMeyerAI.com

You May Also Like

AI On A Budget: The Secret Weapon In The Open-Weight Price War

Alibaba releases a low-cost, capable open-weight AI model to dominate the global developer market amid a fierce price war in the AI industry.

AI Milestone: What DeepSeek-V4-Flash-High Shows At $0.25 Per Million

DeepSeek-V4-Flash-High, a new AI model, is rated highly on Arena’s leaderboard at a cost of approximately $0.25 per million tokens, marking a significant milestone.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs and reorg are framed around AI, but evidence suggests market pressures and cost-cutting are primary. Here’s what’s confirmed and what remains unclear.

Customer service + BPO. The operational-scale displacement.

Empirical evidence shows 8 million workers in India and the Philippines face AI-driven displacement, with a shift to hybrid models as the new operational norm.