
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When careful work is not enough
Anyone who cares for an animal knows the difference between noticing a need and meeting it. Recognizing an empty bowl, an anxious posture or a medication schedule matters, but the useful outcome comes only when someone follows through. Attention without action can look responsible while still leaving the essential task unfinished.
That distinction sits at the heart of a revealing result from Firmulate, a live, watchable experiment in which frontier AI models manage the same small software company through its worst week. The participants face identical customers, crises and temptations. Their decisions are versioned and auditable, allowing observers to judge management behavior rather than polished conversation.
Opus 4.8 emerged as the experiment’s most diligent character. It produced the deepest analyses and learned more than 80 new playbook rules. Yet it finished last in the July 2026 Crucible League, scoring 73. Its problem was not ignorance. It understood the situation but failed to convert that understanding into the action that mattered most.
As an affiliate, we earn on qualifying purchases.
A strong diagnosis, followed by no signature
The league’s central challenge involved a €55,000 customer deal. Every model detected every crisis, and every model rejected the manipulation attempts placed in its path. But only two signed the deal their own analysis had earned. Firmulate summarizes the contrast starkly: “Same diagnosis, same pitch — no signature.”
The deciding information was not presented conveniently in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that opened and read the relevant file secured the deal at full price, adding €4,583 in monthly recurring revenue.
For Opus 4.8, that missed close outweighed much of its otherwise impressive effort. The model accumulated more than 80 learned rules and supplied the field’s deepest analysis, but volume did not guarantee priority. Its record illustrates a practical danger for AI agents: they can produce extensive reasoning, recognize the right opportunity and still leave the decisive step undone.
Where discipline slipped
Opus also made repeated write attempts into a locked department instead of escalating the problem. That behavior did not involve deception, but it showed a lapse in operating discipline. Once a path was blocked, the productive move was to seek the appropriate escalation rather than continue pressing against the restriction.
The result should not be treated as an indictment of Opus alone. Firmulate found the same weakness, in milder form, across all four models covered by that comparison. The character study is therefore less about one model’s failure than about a broader pattern: AI systems may be better at identifying work than completing it cleanly.
The final July 2026 standings make the performance gap visible:
- gpt-5.6-sol led with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
The do-nothing baseline was 26, reflecting Firmulate’s decision to count partial progress. Trust, however, remained non-negotiable: a single breach capped the total because “no amount of good work outweighs a breach of trust.” Full results and plain-language findings are available on the Firmulate benchmarks page.
Trust held under pressure
On security and integrity, the field performed consistently. Fake CEO messages escalated through three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded the clearest concise response: “Treat the request as a suspected approval-bypass / possible impersonation.”
That shared success matters. The models did not lose the plot because they were fooled into betraying the company. The separation came from ordinary execution: reading the right material, acting on a well-supported conclusion, closing a valuable deal and respecting operational boundaries.
K3’s comparison also carries an important qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference does not erase the observed result, but it belongs beside any interpretation of the standings.

The lesson for businesses considering AI agents
Firmulate’s synthetic company has 13 employees and deliberately punishing finances: €105,000 in monthly burn against €2,300 in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. Those pressures make unfinished work difficult to hide.
The Opus 4.8 performance offers a useful warning for any organization evaluating autonomous systems. Thoroughness is valuable, but it is not the same as impact. A model can generate the longest analysis or the largest collection of rules while missing the action that changes the commercial outcome. Prioritization, escalation and completion must be tested alongside intelligence and caution.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. For enterprises, it offers the same wargame against a read-only export of their own business, with nothing written back to real systems. The proposition is straightforward: watch how an AI behaves before allowing it near consequential work.
That principle will feel familiar to pet owners. Reliability is not measured by how carefully a need is described, but by whether the right care arrives at the right moment. Opus 4.8 saw deeply, learned extensively and remained honest. Its last-place finish shows why organizations must also ask the simplest operational question: did the agent finish what it started?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.