
Style
Every fashion editor knows the type: the stylist who glances at you and reaches for whatever’s on the front rack, versus the one who pulls your file, remembers you’re petite, hates high necks, and has a March event coming. The difference isn’t talent. It’s homework. One sells you something; the other actually dresses you.
A new public experiment by Firmulate, which runs AI models as complete simulated companies, suggests the exact same divide exists among frontier AI agents — and it can be measured in euros.
As an affiliate, we earn on qualifying purchases.
The worst week in business, on repeat
Firmulate handed four frontier AI models the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, and the whole thing runs live, watchable at firmulate.com/live, complete with 13 synthetic employees, a public cash countdown, and real money mechanics: €105k a month in burn against just €2.3k in monthly recurring revenue.
By the final July 2026 league table, the standings were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the whole total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The buried fact
Here’s where it gets interesting. Every model spotted every crisis. Every model refused every manipulation attempt. But only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
What separated the closers from the choke artists? A single decisive fact about a competitor’s weakness — buried not in the customer meeting, but two document references deep in the company’s own internal files. The models that actually read the file before answering won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t lost it automatically.
In fashion terms: the stylist who read the client file got the commission. The one who improvised from the front rack watched the sale walk out the door.
As an affiliate, we earn on qualifying purchases.
Honesty under pressure
The experiment also tested nerve. Fake CEO messages escalated over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval-bypass, possible impersonation.
Even the last-place finisher, Opus 4.8, had a fascinating profile: it was the most thorough participant of the field, with the deepest analyses and 80-plus self-learned rules — and still finished last, because the close was left on the table and discipline slipped, including write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
One fairness note: Kimi K3 ran without an effort parameter, at the API default, while its rivals ran at maximum effort — and still placed second.

As an affiliate, we earn on qualifying purchases.
Why it matters
If AI agents are going to touch your CRM, your support queue, or your forecast, the question is no longer “does it write well?” Chat demos can’t show you the gap between a model that diagnoses and a model that finishes. “Reads your files before answering” turns out to be a measurable, purchase-deciding property — worth €4,583 a month in one simulated week.
You can test your own instincts: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html, contact@firmulate.com).
The lesson translates cleanly from the runway to the boardroom: preparation isn’t a personality trait. It’s a feature. And now it’s a line item you can shop for.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html