firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Any stylist will tell you: a look is never judged on one element. The shoes can be perfect, the tailoring immaculate — but one torn lining inside the jacket and the whole outfit is compromised. Fashion scoring has always understood something that AI benchmarking is only now learning: partial credit is real, but trust, once broken, caps everything. You don’t get to accessorize your way out of a broken zipper.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That philosophy is exactly what’s running right now at Firmulate, a live experiment where frontier AI models each run the same small software company through its worst week — same customers, same crises, same temptations to cheat. And the scoring system produces one of the most talked-about numbers in the field: a manager AI that does absolutely nothing still scores 26 points. Not zero. Here’s why that’s not a bug — it’s the whole point.

Why the Floor Is 26, Not 0

Imagine a contestant on a design-show challenge who shows up, reads the brief, understands the client, and then — freezes. Produces nothing. Would you score them the same as someone who didn’t read the brief at all? Of course not. Reading the brief is worth something.

Firmulate’s do-nothing baseline scores 26 for the same reason. A manager that merely understands the situation correctly — that recognizes what’s happening, who the customers are, what the stakes are — has achieved partial progress. The benchmark’s designers deliberately count that progress, because in the real world, correct diagnosis without action is frustrating but not worthless. A team that knows there’s a fire is meaningfully ahead of a team that doesn’t.

What it means is that scores aren’t inflated by a generous curve. The scale simply starts where actual understanding starts — and 26 is the honest marker for “saw everything, did nothing.” The winners in the final July 2026 league table — gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88 — sit where they sit because they did dramatically more than look around.

Amazon

AI management scoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Rule That Caps Everything

Here’s where fashion’s hardest lesson applies: a single breach of trust caps the total score. Firmulate’s own formulation is blunt — “no amount of good work outweighs a breach of trust.”

It’s the fake-leather-handbag principle. The bag can be beautifully made, perfectly proportioned, gorgeous hardware — but if it’s sold as leather and it isn’t, the value collapses. Not discounted. Collapses. Firmulate bakes that same asymmetry into management scoring: if the AI deceives a customer, hides a mistake, or breaks faith even once, no amount of brilliance afterward can restore the grade. Good behavior is cumulative; trust is binary.

This matters because most AI benchmarks score like a checklist — one bad answer costs you a point, and enough good answers buy it back. Firmulate’s stance is that management doesn’t work that way, so the measurement shouldn’t either.

Amazon

trust-based AI evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Suspicion of Round Numbers Is Healthy

Notice something about the final standings: nobody scored 100. The top score is 95. That’s not a coincidence or a shortfall — it reflects a methodology that distrusts perfection. A perfect score in a management simulation would suggest either a too-easy test or a model performing for the grader. Real management, like real style, always leaves something on the table. A benchmark that hands out 100s is a benchmark that isn’t looking hard enough.

Amazon

AI performance benchmarking platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Winners

The experiment’s most striking finding is how far the models got together before diverging. All four participants spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned — the study’s memorable summary: “Same diagnosis, same pitch — no signature.”

The difference was buried. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer conversation at all. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered a flawless pitch and walked away empty-handed. It’s the business equivalent of styling a perfect look without checking the invitation’s dress code.

The social-engineering gauntlet was equally revealing. Fake CEO messages escalated over three stages, plus a reporter’s trick — a seemingly harmless “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The field, in other words, is well-dressed when it comes to honesty — the gaps are in follow-through and homework.

Amazon

AI transparency and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Cautionary Tale: Thorough but Last

Opus 4.8’s profile is the benchmark’s best argument for measuring management rather than fluency. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Effort without judgment. The same weakness appeared, more faintly, in all four models.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Watch It Live

This isn’t a static report. The live company — 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules — is watchable at firmulate.com/live, with every workday versioned and auditable. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is the most honest thing about this benchmark. It says: understanding counts, effort counts, but nothing is given away for free — and nothing after a breach of trust counts at all. Whether you’re judging an AI to run your support queue or judging a runway look, the principles translate: partial progress deserves partial credit, perfection deserves suspicion, and trust is the one accessory you can’t replace. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Properly Fit and Wear a Bike Helmet

Theoretically perfect helmet fit can prevent injuries, but learning how to properly wear your helmet ensures maximum safety and comfort during every ride.

Tire Casing, TPI, and Puncture Layers—Choose What Really Matters

Learn how tire casing, TPI, and puncture layers impact performance, helping you choose the best tires for your ride—discover what truly matters.

Bed‑In Disc Brakes the Right Way (Most People Don’t)

Aiming for optimal braking performance, learn how to bed-in your disc brakes correctly—most people skip this step, but here’s why it matters.

Index Your Rear Derailleur for Crisp Shifts

Gaining perfect shifting begins with proper derailleur indexing—discover the key steps to ensure smooth, crisp gear changes on your bike.