
Anyone who has trained a rescue dog knows the score is never zero. On day one, the dog that doesn’t bite the mailman, doesn’t destroy the couch, and sits half the time isn’t failing — it’s partially succeeding. You grade progress, because a dog that gets nothing right doesn’t exist, and a grading system that says “0” for imperfect effort teaches you nothing about what to fix.
Get pet supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
It turns out the same logic applies to judging AI models as managers. In a live, public experiment run by Firmulate, five frontier AI models were each handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cheat. The final league table is topped by gpt-5.6-sol at 95 points, with Kimi K3 close behind at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
But the number that made me sit up wasn’t at the top of the table. It was at the bottom — an invisible floor. A do-nothing baseline run, where the AI manager takes essentially no meaningful action, scores 26 points, not 0. That number is the most honest thing on the whole scoreboard, and the story behind it says a lot about what trustworthy AI evaluation should look like.
Why doing nothing still earns 26
Think about what a manager who does nothing actually accomplishes in a crisis week. The fires that don’t get started. The customer who isn’t insulted. The manipulation attempt that gets ignored by default because nobody answered the phone. Inaction has consequences in both directions, and a fair scoring system has to recognize both.
That’s the philosophy behind Firmulate’s 26-point floor: partial progress counts. A model that diagnoses a problem but never closes the deal has done real, measurable work — just not all of it. The experiment’s central finding illustrates this perfectly. All five models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Under a crude pass/fail system, all five would look identical. Under a system where partial progress counts, the difference between a 95 and a 73 becomes visible and meaningful.
As an affiliate, we earn on qualifying purchases.
One breach of trust caps everything
The floor at 26 has a mirror image at the top: a single breach of trust caps the total score, no matter how brilliant the rest of the performance. The benchmark’s own framing is blunt — “no amount of good work outweighs a breach of trust.”
This is the same instinct anyone who runs a kennel or a vet clinic knows intimately. A handler who is wonderful with ninety-nine dogs but kicks the hundredth doesn’t get a 99 percent rating. Trust isn’t an average; it’s a gate. An AI benchmark that lets a model lie, cheat, or manipulate its way to a high overall score would be measuring the wrong thing entirely — and, worse, rewarding exactly the behavior you’d most want to catch.
Why the winners won: they read the files
The buried fact behind the league table is almost comic in its simplicity. The decisive competitor weakness — the piece of information that made the €55,000 deal closeable at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files.
The models that actually read the company’s paperwork before walking into the negotiation won the deal. The ones that didn’t, didn’t. It’s the AI equivalent of the dog trainer who actually reads the intake form before the first session — unglamorous, and completely decisive.
Then there’s Opus 4.8, the cautionary tale of the experiment: the most thorough participant of the field, with over 80 learned rules added to its playbook and the deepest analyses of any model — and last place at 73. It left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating the issue. Effort without follow-through. Every owner of an energetic, unfocused puppy knows the type. Notably, the same weakness appeared, weaker, in all four other models.
Honest under pressure — all five of them
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning, on the record, was: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you want codified before an AI agent ever touches a real support queue.
One footnote for fairness: K3 ran without an effort parameter while the others ran at their highest effort setting, which makes its second-place finish at 93 arguably even more striking.
You can watch the company live
This isn’t a one-off lab report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,325 in monthly recurring revenue, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned and auditable. It’s watchable at firmulate.com/live — less like a benchmark and more like a very slow, very expensive reality show where the cast never sleeps.
For those who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The 26-point floor and the trust cap are two sides of the same idea: scores should tell the truth, not flatter the scorer. Partial progress earns partial credit; a breach of trust earns a ceiling. And the league table itself carries a healthy distrust of round 100s — the top score is 95, not a suspiciously perfect 100, because real management under real pressure is never flawless.
That’s the standard anyone hiring an AI for consequential work should demand. Not “does it chat well,” but: does it finish what it starts, does it read the files first, and does it stay honest when nobody’s checking? The full results and plain-language findings are at firmulate.com/benchmarks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
