
Style Isn’t the Same as Substance — for Models Too
Anyone who has worked in fashion knows the type: the flawless look, the perfect pitch, the front row seat — and then the sample that never ships. Presentation is not the same as performance. It turns out artificial intelligence has exactly the same problem, and someone finally ran the experiment to prove it.
While the tech world crowns its AI champions on coding leaderboards and chat arenas — contests that measure how elegantly a model answers — a live experiment at Firmulate asked a different question: what happens when an AI has to actually manage? Run a company. Handle a churn wave. Survive a PR crisis. Close a deal. Stay honest when a fake CEO tries to talk it into something stupid.
The results read like a season finale: all four frontier models looked brilliant, and only some of them delivered.
As an affiliate, we earn on qualifying purchases.
One Company, Four CEOs, the Worst Week Ever
The setup is elegantly simple, the way the best runway concepts are. Four frontier AI models — GPT-5.6-Sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 among the field — were each handed the same small software company and the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retried.
The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 last at 73. For context, doing absolutely nothing scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.
(One fairness footnote worth flagging, in the spirit of honest labeling: Kimi K3 ran at its API-default effort setting while the others ran at xhigh. It placed second anyway.)
AI customer relationship management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Finding That Chat Demos Can’t Show
Here’s what should worry anyone planning to put an AI agent near their CRM, support queue or forecast: every single model spotted every crisis, and every single one refused every manipulation attempt. Five out of five models saw through a fake-CEO social-engineering campaign that escalated over three stages, plus a reporter’s “just one yes/no, on background” trick. Kimi K3 said it on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Sharp, composed, unfoolable.
And yet only two models signed the €55,000 deal that their own analysis had earned them. Same diagnosis, same pitch — no signature. They dressed for the meeting and then left before the contract. It’s the AI equivalent of a collection that photographs beautifully and can’t survive a wash cycle.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
The most instructive detail of the whole experiment is almost invisible. The decisive competitive weakness — the fact that should have powered the winning pitch — wasn’t in the customer call or the crisis inbox. It sat two document references deep in the company’s own files. The models that actually read their own paperwork found it and closed the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
Preparation beats improvisation. Some things never change, whether you’re a buyer preparing for market week or an AI preparing for a sales call.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hardest Working Model Came Last
Then there’s Opus 4.8, the cautionary tale of the season. It was the most thorough participant in the entire field — over 80 self-learned playbook rules, the deepest analyses of anyone. And it finished dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness showed up, more faintly, in all four models. Effort without judgment is just expensive activity.
You Can Watch the Company Lose Money in Real Time
This isn’t a slide deck or a press release. The company at the heart of the experiment is live software with 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com.
Want to test your own eye? There’s a quiz built from 242 real, unedited management decisions where you guess which model made which call. Full results and plain-language findings live on the benchmarks page. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Management Quality, Not Chat Quality
The lesson generalizes far beyond software companies. The benchmarks that currently decide which AI is “best” measure how well a model converses — the equivalent of judging a designer by their mood board. Firmulate’s experiment measures what happens after the conversation: does the agent finish what it starts, does it read the files first, does it stay honest under pressure, and what does a unit of useful work actually cost?
The scenario names — churn wave, price increase, downround, PR crisis — are the new curriculum. And the early evidence says the gap between the two kinds of quality is wide: a field of models that aced every conversation and refused every trap, where barely half managed to sign the deal sitting in front of them.
Before you hire the best-dressed AI in the room, watch it work a full week. The ones that look good in the demo and the ones that deliver are, it turns out, not always the same models — and now there’s a scoreboard that finally shows the difference.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html