AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Knowing the answer is not the same as taking responsibility

Anyone who cares for animals understands the distance between recognizing a problem and managing it well. Spotting an open gate, a missed meal or an anxious pet matters, but the real test is what happens next: Does someone act, check the details and follow through without creating another risk?

Artificial intelligence faces a similar gap. Coding leaderboards and chat arenas can tell us whether a model produces an impressive answer. They reveal much less about whether an AI agent can prioritize competing demands, withstand pressure, protect confidential information and complete commercially important work across days.

Firmulate, a live AI company experiment, is trying to measure that missing capability. Its proposition is timely: businesses preparing to give agents access to customer relationships, support queues or forecasts should evaluate management quality, not merely chat quality.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations remained constant, and every decision was versioned and auditable. The scenarios included a churn wave, a price increase, a downround and a PR crisis—the sort of situations in which consequences accumulate and yesterday’s shortcut becomes tomorrow’s emergency.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But trust is treated as a hard boundary: “no amount of good work outweighs a breach of trust.”

Those results tell a more interesting story than a conventional ranking. All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary is brutally concise: “Same diagnosis, same pitch — no signature.”

This is the measurement gap in miniature. A model can understand the situation, draft the right material and still fail to produce the outcome. In a chat window, a polished response looks complete. Inside a company, completion may require reading another document, escalating an obstacle or making the final move.

The decisive fact was not where the action happened

The detail that separated the strongest performances was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that found it won the deal at full price, worth +€4,583 MRR.

That finding should matter to any organization considering AI agents. Business work rarely arrives as a self-contained prompt. The crucial context may sit in a contract, an earlier decision or a neglected internal record. Fluent improvisation cannot substitute for looking in the right place before acting.

Opus 4.8 makes the point especially well. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

Thoroughness, then, is not synonymous with management. A capable manager must convert analysis into action while respecting boundaries. More reasoning can coexist with a missed outcome.

Pressure tested honesty as well as competence

The experiment also subjected the models to fake CEO messages that escalated over three stages, plus a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is reassuring, but it also shows why isolated answer-quality tests are insufficient. Honesty becomes meaningful when disclosure is tempting, authority appears plausible and the company is already under strain. Firmulate tests whether models preserve judgment amid those conditions.

There is an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when readers interpret the narrow gap at the top of the table.

A company that can be watched

This is not presented as a fictional management exercise. The live company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Readers can follow the experiment through Firmulate and examine the final results on its public benchmark page.

The project also turns 242 real, unedited management decisions into a “guess the model” quiz. For enterprises, its pilot offers the same kind of wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The next benchmark should ask who can be trusted with the keys

For businesses, the lesson is not that coding tests or chat evaluations are useless. It is that they measure only part of the job. Agents operating inside companies must discover context, triage under capacity pressure, finish valuable work and remain honest when manipulation arrives wearing an executive’s name.

Scenario names such as churn wave, price increase, downround and PR crisis point toward a new curriculum for AI evaluation. The central question is no longer simply whether a model can produce the right answer. It is whether the model can manage the whole situation—and close the gate after noticing that it is open.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Pet-care content is informational — consult your veterinarian for advice about your animal.


You May Also Like

Proofing for Trials: Distractions, Surfaces, and Weather

Staying focused during proofing for trials requires managing distractions, surfaces, and weather—discover essential tips to stay prepared and confident in any environment.

Reward Placement Changes Everything in Agility

Meta description: “Mastering reward placement can unlock your dog’s agility potential—discover how subtle changes may be the key to endless progress.

Advanced Agility Drills: Boosting Your Dog’s Speed and Precision

Cleverly incorporating advanced agility drills can significantly enhance your dog’s speed and precision, but mastering these techniques requires guidance and consistent practice.

Small Space Agility: Drills for Apartments

Tiny apartment spaces can still boost agility with creative drills; discover how to transform your limited area into an effective workout zone.