firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Anyone who follows fashion knows the story: an unknown label arrives, everyone expects a knock-off, and then the collection walks off with the season. It happened with the Japanese designers who shook Paris in the eighties, and it happens every few years with a debut line that simply executes better than the establishment. This month, the same plot played out — not on a runway, but in a strange and very public experiment where AI models were put in charge of running a small software company through its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment, run by Firmulate, gave four frontier AI models the identical job: run the same company, face the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable. A fifth entrant — Kimi K3, made by Moonshot — joined the field, and the final league table reads like a debut collection beating the established houses.

The standings: gpt-5.6-sol took first with 95. Kimi K3, the newcomer, scored 93 — second place, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. For context, a do-nothing baseline scores 26, and a single breach of trust caps a model’s total outright — in the experiment’s own words, “no amount of good work outweighs a breach of trust.”

What actually happened in that week

The pattern across the field was striking: all models spotted every crisis and refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. The experiment’s summary of the others: “Same diagnosis, same pitch — no signature.” That deal was worth an additional €4,583 in monthly recurring revenue.

The deal turned on something buried: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file won the deal at full price. The ones that didn’t, didn’t.

K3 also found a buried security needle, saved a churning customer, and resisted all three social-engineering baits — including a reporter’s “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, K3 recorded only one deviation — the cleanest discipline in the field.

The cautionary tale at the bottom of the table

The most fascinating result belongs to Opus 4.8: the most thorough participant of all, generating more than 80 learned rules and the deepest analyses — and still last place. It left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

It’s the AI equivalent of a collection with impeccable construction and no edit: all craft, no finish.

Why this matters beyond the league table

The lesson for anyone choosing an AI tool — for a styling business, a shop floor, a support queue — is the same lesson fashion teaches about hype: the established label doesn’t always win on execution. The league is open. Picking a model without testing it against your own work is now a bet, not a decision.

You can watch the underlying company live at firmulate.com — 13 synthetic employees, real money mechanics (a €105k monthly burn against €2.3k in MRR, with a public cash countdown), and more than 680 self-learned playbook rules, versioned every workday. Full results are on the benchmarks page, a “guess the model” quiz draws on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.

A note on fairness: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The debut line beat three of four established houses — not by being flashier, but by finishing: reading the file, closing the deal, staying clean under pressure. Judge the collection, not the label. Before you commit to any AI model, run it through your own worst week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI customer service automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise risk management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Digital Floor Pumps: Do They Really Make Inflation Easier?

Just how much do digital floor pumps simplify inflation, and can they truly enhance your experience? Find out more to discover the answer.

Helmet Fit That Stays Put: The Two‑Finger Test, Upgraded

Ineffective helmet fit can compromise safety—discover the upgraded two-finger test to ensure your helmet stays secure and reliable.

Setting Up Your Bike for a Smooth Ride

Navigate your bike setup for a smooth ride; discover essential tips that can transform your cycling experience into something truly enjoyable.

Mini Pumps vs Electric Air Pumps: Which One Fits Your Routine?

Finding the perfect pump depends on your needs—whether portability or efficiency—and exploring options will help you make the best choice.