AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Can you read an AI manager as well as you read your pet?

Pet owners become fluent in behavior. A dog hovering near the door, a cat abruptly refusing its usual perch or a bird going quiet can communicate more than an elaborate display. The interesting question is not whether the animal looks clever, but what it consistently does when something matters.

Firmulate applies a similar behavioral lens to frontier artificial intelligence. Instead of judging models by polished conversation, it placed them in charge of the same small software company during its worst week. They faced identical customers, crises and temptations. Their decisions were preserved and could be audited, turning management style into observable evidence rather than marketing copy.

Those decisions now form an unusually revealing reader challenge. The Firmulate quiz draws on 242 real, unedited management decisions and asks players to identify which model made each one. The differences are not cosmetic. One participant produced exceptionally deep analysis, others completed the commercial task, and another explicitly treated a suspicious instruction as a possible impersonation attempt.

Amazon

AI management decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical problems, distinctly different behavior

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline earned 26 because partial progress counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”

Every model identified every crisis and rejected every manipulation attempt. That shared competence makes the commercial result more striking. Only two models signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The decisive information was not sitting prominently in a customer event. It was buried two document references deep inside the company’s own files. The models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. The episode illustrates a practical distinction for any business considering AI workers: noticing a problem is not the same as gathering the right evidence, acting on it and completing the transaction.

Pressure exposed discipline as well as intelligence

The models also encountered fake messages from the chief executive that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because apparent helpfulness can become a liability when an instruction is deceptive. In this experiment, the entire field maintained the boundary. The more revealing separation came from operational follow-through: reading deeply enough, respecting constraints and finishing valuable work.

Opus 4.8 makes the point especially well. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the commercial close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

There is also an important fairness qualification. Kimi K3 ran with the application programming interface’s default setting because it had no effort parameter, while the other models ran at xhigh. That does not erase its result, but it belongs beside the ranking so readers can interpret the comparison responsibly.

A company designed to make consequences visible

Firmulate’s live company employs 13 synthetic workers and uses real money mechanics. It burns €105,000 each month while producing €2,300 in monthly recurring revenue, with a public cash countdown making the imbalance visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

This is why the quiz feels less like trivia and more like field observation. A terse response may belong to a model that acts decisively. An exhaustive memorandum may come from one that understands everything but fails to close. A refusal may represent sound judgment rather than timidity. Readers are being asked to identify durable management behavior from the evidence left behind.

Infographic —
The findings at a glance — source: firmulate.com.

The personality is in the follow-through

For pet-focused readers, the central lesson will feel familiar: character becomes clearest through repeated behavior under pressure. The most expressive performance is not necessarily the most dependable, and apparent intelligence is only part of the relationship.

Firmulate’s experiment shows that frontier models can agree about risks while differing materially in execution. They may all spot a crisis and resist manipulation, yet separate when success depends on reading overlooked files, maintaining discipline and completing a deal.

The quiz makes those distinctions accessible without asking readers to decode technical benchmarks. It presents the decisions themselves and lets the audience form a judgment before seeing which model was responsible. In a market crowded with impressive demonstrations, that is a useful reversal: watch what the AI manager actually does, then decide what kind of manager it is.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Pet-care content is informational — consult your veterinarian for advice about your animal.


You May Also Like

Agility Warmups Most Handlers Skip

Great warmup routines can prevent injuries and boost performance, but many handlers overlook essential steps that could make a big difference for your dog.

Weave Pole Entries: The Tiny Detail That Makes Speed Reliable

Lacking small but crucial details in weave pole entries can hinder your dog’s speed; discover the subtle adjustments that make entries faster and more reliable.

Flatwork First: The Agility Skill Hidden in Plain Sight

Absolutely essential yet often overlooked, flatwork first unlocks hidden agility skills that can elevate your precision and safety—discover how inside.

Agility Training for Dogs: Mental and Physical Benefits

Overcome your dog’s physical and mental challenges with agility training; discover how it can transform their confidence and strengthen your bond.