firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Anyone who follows fashion knows the story: an unknown label arrives, everyone expects a knock-off, and then the collection walks off with the season. It happened with the Japanese designers who shook Paris in the eighties, and it happens every few years with a debut line that simply executes better than the establishment. This month, the same plot played out — not on a runway, but in a strange and very public experiment where AI models were put in charge of running a small software company through its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment, run by Firmulate, gave four frontier AI models the identical job: run the same company, face the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable. A fifth entrant — Kimi K3, made by Moonshot — joined the field, and the final league table reads like a debut collection beating the established houses.

The standings: gpt-5.6-sol took first with 95. Kimi K3, the newcomer, scored 93 — second place, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. For context, a do-nothing baseline scores 26, and a single breach of trust caps a model’s total outright — in the experiment’s own words, “no amount of good work outweighs a breach of trust.”

What actually happened in that week

The pattern across the field was striking: all models spotted every crisis and refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. The experiment’s summary of the others: “Same diagnosis, same pitch — no signature.” That deal was worth an additional €4,583 in monthly recurring revenue.

The deal turned on something buried: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file won the deal at full price. The ones that didn’t, didn’t.

K3 also found a buried security needle, saved a churning customer, and resisted all three social-engineering baits — including a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, K3 recorded only one deviation — the cleanest discipline in the field.

The cautionary tale at the bottom of the table

The most fascinating result belongs to Opus 4.8: the most thorough participant of all, generating more than 80 learned rules and the deepest analyses — and still last place. It left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

It’s the AI equivalent of a collection with impeccable construction and no edit: all craft, no finish.

Why this matters beyond the league table

The lesson for anyone choosing an AI tool — for a styling business, a shop floor, a support queue — is the same lesson fashion teaches about hype: the established label doesn’t always win on execution. The league is open. Picking a model without testing it against your own work is now a bet, not a decision.

You can watch the underlying company live at firmulate.com — 13 synthetic employees, real money mechanics (a €105k monthly burn against €2.3k in MRR, with a public cash countdown), and more than 680 self-learned playbook rules, versioned every workday. Full results are on the benchmarks page, a “guess the model” quiz draws on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.

A note on fairness: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The debut line beat three of four established houses — not by being flashier, but by finishing: reading the file, closing the deal, staying clean under pressure. Judge the collection, not the label. Before you commit to any AI model, run it through your own worst week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI customer service automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise risk management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Direct Drive vs. Wheel-On Trainers: Pros, Cons, and What You Need to Know

Make an informed choice between direct drive and wheel-on trainers to elevate your cycling experience—discover which option suits your needs best!

Physics-based Interaction: A Look Inside “The Belfry | Eight Bells, Endless Changes” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Belfry…

Winterize Your Bike: Salt, Lube, and Protection

Learn how to winterize your bike with salt protection, proper lubrication, and essential maintenance to keep it in top shape all season long.

When the Boss Demands a Shortcut, the Smartest AI Says No

Five frontier AI models rejected fake CEO pressure and a reporter’s trick, showing companies can test integrity before agents reach production.