firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Style Isn’t the Same as Substance — for Models Too

Anyone who has worked in fashion knows the type: the flawless look, the perfect pitch, the front row seat — and then the sample that never ships. Presentation is not the same as performance. It turns out artificial intelligence has exactly the same problem, and someone finally ran the experiment to prove it.

While the tech world crowns its AI champions on coding leaderboards and chat arenas — contests that measure how elegantly a model answers — a live experiment at Firmulate asked a different question: what happens when an AI has to actually manage? Run a company. Handle a churn wave. Survive a PR crisis. Close a deal. Stay honest when a fake CEO tries to talk it into something stupid.

The results read like a season finale: all four frontier models looked brilliant, and only some of them delivered.

Amazon

AI contract signing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Four CEOs, the Worst Week Ever

The setup is elegantly simple, the way the best runway concepts are. Four frontier AI models — GPT-5.6-Sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 among the field — were each handed the same small software company and the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retried.

The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 last at 73. For context, doing absolutely nothing scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.

(One fairness footnote worth flagging, in the spirit of honest labeling: Kimi K3 ran at its API-default effort setting while the others ran at xhigh. It placed second anyway.)

Amazon

AI customer relationship management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Can’t Show

Here’s what should worry anyone planning to put an AI agent near their CRM, support queue or forecast: every single model spotted every crisis, and every single one refused every manipulation attempt. Five out of five models saw through a fake-CEO social-engineering campaign that escalated over three stages, plus a reporter’s “just one yes/no, on background” trick. Kimi K3 said it on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Sharp, composed, unfoolable.

And yet only two models signed the €55,000 deal that their own analysis had earned them. Same diagnosis, same pitch — no signature. They dressed for the meeting and then left before the contract. It’s the AI equivalent of a collection that photographs beautifully and can’t survive a wash cycle.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The most instructive detail of the whole experiment is almost invisible. The decisive competitive weakness — the fact that should have powered the winning pitch — wasn’t in the customer call or the crisis inbox. It sat two document references deep in the company’s own files. The models that actually read their own paperwork found it and closed the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Preparation beats improvisation. Some things never change, whether you’re a buyer preparing for market week or an AI preparing for a sales call.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hardest Working Model Came Last

Then there’s Opus 4.8, the cautionary tale of the season. It was the most thorough participant in the entire field — over 80 self-learned playbook rules, the deepest analyses of anyone. And it finished dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness showed up, more faintly, in all four models. Effort without judgment is just expensive activity.

You Can Watch the Company Lose Money in Real Time

This isn’t a slide deck or a press release. The company at the heart of the experiment is live software with 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com.

Want to test your own eye? There’s a quiz built from 242 real, unedited management decisions where you guess which model made which call. Full results and plain-language findings live on the benchmarks page. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The lesson generalizes far beyond software companies. The benchmarks that currently decide which AI is “best” measure how well a model converses — the equivalent of judging a designer by their mood board. Firmulate’s experiment measures what happens after the conversation: does the agent finish what it starts, does it read the files first, does it stay honest under pressure, and what does a unit of useful work actually cost?

The scenario names — churn wave, price increase, downround, PR crisis — are the new curriculum. And the early evidence says the gap between the two kinds of quality is wide: a field of models that aced every conversation and refused every trap, where barely half managed to sign the deal sitting in front of them.

Before you hire the best-dressed AI in the room, watch it work a full week. The ones that look good in the demo and the ones that deliver are, it turns out, not always the same models — and now there’s a scoreboard that finally shows the difference.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Discover practical tips on acoustic dampening, placement, and the ‘rig in the closet’ trick to make your workspace quieter and sound better — without breaking the bank.

Upgrading Your Bike on a Budget: Tips and Tricks

Never settle for a basic bike; discover budget-friendly upgrades that will transform your ride and enhance your cycling experience in surprising ways.

How to Maintain Your Bike: A Comprehensive Beginner’s Guide

Learn essential bike maintenance tips and tricks to keep your ride in top shape; discover what crucial tools every beginner should have.

Store an E‑Bike Battery Safely: Charge Levels and Temps

Just knowing the right charge levels and temperatures can significantly extend your e-bike battery’s lifespan—discover how to store it safely.