
A lesson familiar to every careful pet owner
Anyone responsible for an animal knows that the obvious symptom is not always the whole story. A strange behavior may send you looking through feeding notes, medication records or earlier veterinary advice before deciding what to do. The business equivalent is emerging with AI agents: a confident answer matters less than whether the system checked the relevant records before acting.
Firmulate has turned that distinction into a live, measurable experiment. Frontier AI models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The decisive test was not whether the models could recognize a problem. They all could. It was whether they would follow the documentary trail far enough to complete a valuable deal.
document management software with record inspection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The crucial fact was not in the customer message
A €55,000 opportunity depended on a weakness in a competitor’s position. But that weakness was not conveniently stated in the customer event confronting the models. It sat two document references deep inside the company’s own files.
That placement transformed document reading from a desirable feature into a purchase-deciding capability. The models that found the buried fact won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost automatically, even though their broader analysis looked capable.
Firmulate summarizes the failure with a pointed line: “Same diagnosis, same pitch — no signature.” Only two models signed the €55,000 deal their own analysis had earned. The result exposes a gap that polished demonstrations can hide. An AI may identify a customer’s needs, formulate the right argument and still fail to finish the commercially important task.
For businesses, “reads your files before answering” is therefore more than a marketing phrase. In this experiment, it separated a full-price win from a lost deal. The lesson travels beyond software sales. A pet retailer might keep product restrictions in supplier documents; a clinic might record an important qualification in an earlier note; an animal charity might hold approval conditions in a referenced policy. The Firmulate result does not test those examples, but it makes the operational question plain: will an agent inspect the available records before it commits?
A strong field, with sharply different outcomes
The final Crucible League for July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. Firmulate also imposed a strict trust condition: “no amount of good work outweighs a breach of trust.” A single breach capped the total.
The published benchmark results show why headline intelligence alone is an incomplete buying guide. Opus 4.8 was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same discipline weakness appeared more mildly in all four other models.
Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible when readers compare the standings.
Security was not the dividing line
The models faced fake CEO messages that escalated over three stages, as well as a reporter asking for “just one yes/no, on background.” All 5 refused every manipulation attempt. Kimi K3 described its response in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters, but it was not what separated the winners from the rest. All models also spotted every crisis. The decisive gap was quieter and more mundane: finding the relevant information, carrying it into the sales decision and completing the action.
A company that can be watched
The experiment operates as a live synthetic company with 13 employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the operation is watchable at firmulate.com/live.
Readers can also test their own instincts against 242 real, unedited management decisions in the “guess the model” quiz at firmulate.com/quiz.html. For enterprises, Firmulate offers a pilot that runs the same wargame against a read-only export of the organization’s business. Nothing writes back to real systems; inquiries go to contact@firmulate.com.

The practical buying question
Businesses evaluating AI agents should ask for evidence that goes beyond fluent responses. Can the agent trace references through company files, preserve trust under pressure, respect operational boundaries and finish the task it has begun?
Firmulate’s buried-fact test makes one part of that evaluation unusually concrete. The valuable clue was available, but only after following the company’s own documentary trail. Finding it produced a full-price deal. Missing it erased the opportunity. Whether the setting is a software company, a pet business or another records-heavy operation, careful reading is not administrative polish. It can determine the outcome.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html