firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every fashion person knows someone like this: the friend with the impeccable wardrobe, the ten Pinterest boards, the color-season analysis, the spreadsheet of capsule pieces — who still shows up overdone while the woman in one great blazer and clean shoes walks away with the room. Effort, we learn early in style, is not the same thing as impact. The outfit that wins is rarely the one with the most pieces. It’s the one with priorities.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

It turns out AI models have the same problem — and one of them in spectacular fashion.

The worst week in business, run four times

On the live research platform Firmulate, four frontier AI models were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned afterward. Think of it as a runway show where every model wears the same brief, and the judges score execution, not the lookbook.

The final league table from the July 2026 run: gpt-5.6-sol in first with 95, Kimi K3 at 93, Sonnet 5 at 88, another Sonnet configuration at 77 — and Opus 4.8 in last place at 73. For context, doing nothing at all scores 26, and a single breach of trust caps your total, no matter how much good work came before it.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the vibe check

Here’s the part that should comfort anyone worried about AI ethics, at least in this arena: all four models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, followed by a reporter pulling the classic “just one yes/no, on background” trick — five of five models in the wider study refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the basics are handled. The fashion equivalent: nobody showed up in socks with sandals. The differentiation happened at a much higher level of craft.

Amazon

business crisis management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The deal left on the table

The week included a real prize: a €55,000 deal that the models had to earn themselves. All four diagnosed the customer’s problem correctly. All four made the pitch. Only two signed. Same diagnosis, same pitch — no signature. It’s the business version of a perfect styling consultation where the client walks out without buying anything.

And the reason is almost embarrassingly simple. The decisive weakness in the competing vendor — the fact that justified closing at full price, worth +€4,583 in monthly recurring revenue — wasn’t hidden in the customer conversation. It was buried two document references deep in the company’s own files. The models that actually read the file won the deal. The ones that didn’t, didn’t. Reading before dressing, it turns out, applies to AI too.

Amazon

AI ethics compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Enter Opus 4.8: the overachiever in last place

This is where the story becomes a character study fashion readers will recognize instantly. Opus 4.8 was, by volume, the star of the experiment. It wrote 80 self-learned playbook rules over the course of the run — the most thorough participant by a wide margin, with the deepest analyses of any model in the field. It did the most homework. It built the biggest wardrobe.

And it finished last.

Two things sank it. First, the close: like the other also-rans, it left the €55,000 signature on the table after doing all the analytical work to earn it. Second, discipline slipped under pressure — at one point it made repeated write attempts into a locked department instead of recognizing the lock and escalating the problem to someone who could open it. Persistence in the wrong direction is still the wrong direction.

To be fair, Firmulate’s own findings note that the same weakness appeared, weaker, in all four models. Opus just had the most of it. And one housekeeping note: Kimi K3 ran at the API default effort setting while the others ran at the maximum, which makes its second-place finish arguably the most impressive line on the table.

Amazon

AI reading comprehension tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this matters beyond benchmarks

The live company behind all this is worth watching in its own right: 13 synthetic employees, real money mechanics — a €105k monthly burn against €2.3k in current MRR — a public cash countdown, 680+ self-learned playbook rules, and a full audit trail of every workday. It rebuilds itself twice a day and is viewable at firmulate.com/live.

There’s also a genuinely fun artifact: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a bit like a blind fragrance test, but for management style. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The lesson from Opus 4.8 is the lesson every great stylist eventually internalizes: diligence is not impact. Eighty rules, the deepest analyses, the most effort — and the deal still unsigned, because reading one more file and closing one more loop beats producing three more pages of beautiful work. Prioritization beats volume, for AI agents as much as for wardrobes. If you’re ever asked to bet on an AI colleague, don’t ask which one worked hardest. Ask which one finished.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Silence the Pain Cave: Trainer Noise Fixes That Work

The trainer noise fixes that work can transform your workout space, helping you silence distractions and stay focused—discover how to create a peaceful environment.

DIY Bike Repairs: Fixing Common Issues on the Road

Always be prepared for bike troubles with simple DIY fixes that can keep you rolling; discover essential tips to tackle common road issues!

How to Fix a Flat Tire: A Step-by-Step Guide for Beginners

Nearing a flat tire can be stressful, but this step-by-step guide will teach you everything you need to know to fix it confidently.

Upgrading Your Bike on a Budget: Tips and Tricks

Never settle for a basic bike; discover budget-friendly upgrades that will transform your ride and enhance your cycling experience in surprising ways.