
Management decisions have a signature, too
In fashion, taste reveals itself through choices: what gets emphasized, what gets edited out and whether the final look actually comes together. Artificial intelligence appears to have a comparable tell. Give several frontier models the same company, crises and temptations, and each develops a recognizable management personality.
That premise drives Firmulate’s interactive challenge, built from 242 real, unedited management decisions. Readers see what an AI executive did and try to identify the model behind it. The game is entertaining, but its underlying question is serious: can polished language conceal meaningful differences in judgment, persistence and professional discipline?
AI management decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week at the same company
Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, internal problems and opportunities to take shortcuts remained constant. Every workday and decision was versioned and auditable, turning what could have been a subjective comparison into a watchable business experiment.
The company itself has 13 synthetic employees and deliberately uncomfortable economics: it burns €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, and its workforce has accumulated more than 680 self-learned playbook rules. This is not a test of whether a chatbot can produce a confident memo. It is a test of whether an AI manager can find relevant information, resist manipulation and complete commercially useful work.
The models agreed—until execution mattered
Every model recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the contradiction neatly: “Same diagnosis, same pitch — no signature.” The result suggests that understanding a problem and carrying a decision through to completion are separate management abilities.
The pivotal clue was not delivered in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. In a real organization, that distinction could separate an assistant that sounds informed from an operator that genuinely uses institutional knowledge.
Pressure revealed discipline
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning in unusually direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean sweep matters because management style is not merely tone. A verbose model and a terse model may both reach a safe decision, but their habits become visible in how they investigate, document and finish the work. The Firmulate quiz turns those habits into clues, asking readers to distinguish among decisions without rewriting or polishing them for presentation.
A league table with an unexpected loser
The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
Opus 4.8 produced the most thorough performance, adding +80 learned rules and delivering the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly: intelligent work could still stall when an obstacle demanded a change in approach.
The comparison also carries an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed behavior, but it belongs beside the ranking when readers interpret the result.

As an affiliate, we earn on qualifying purchases.
What the quiz exposes
Fashion teaches us that consistency is more revealing than a single statement piece. The same may be true of AI management. A model’s identity emerges through repeated choices: whether it reads the files, protects trust, adapts when blocked and closes the deal after doing the intellectual work.
Firmulate’s experiment makes those differences tangible. Its quiz offers a playful way into a consequential business question: organizations are not merely choosing writing styles when they select AI agents. They may also be choosing distinct habits of attention, caution and follow-through.
Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That makes the public challenge more than a personality test. It is a preview of how companies might audition an AI workforce before placing it anywhere near customers, forecasts or sensitive operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI management personality assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.