AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What a Couture House Teaches Us About Grading AI

Anyone who has bought a luxury bag knows the unwritten rules. The stitching can be flawless, the leather buttery, the hardware gleaming — but if the authenticity card is fake, the entire piece is worthless. A single counterfeited detail destroys everything. Conversely, a beautifully made piece from a small workshop doesn’t need to be Hermès to have real value: honest materials, honest stitching, honest price.

Now imagine someone graded handbags with the same honesty. A bag that does nothing wrong but also does nothing remarkable wouldn’t score zero — the stitching is real, the leather is genuine, the work is there. It would earn a modest score. But a bag with one forged serial number? Capped, no matter how gorgeous the rest.

That, in essence, is the philosophy behind Firmulate’s AI benchmark — a live experiment that runs frontier AI models as complete companies through their worst possible week and grades them like a maison’s quality inspector: partial progress counts, and a single breach of trust caps the grade. When the final league table published in July 2026, the most talked-about number wasn’t the winner’s 95. It was the floor: a do-nothing baseline scored 26.

Amazon

luxury handbag authentication kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Do-Nothing Baseline: Why 26, Not 0?

Before any AI model got its turn running the fictional software company, Firmulate ran a control: an agent that essentially sat on its hands. It didn’t close deals, didn’t defuse crises, didn’t chase opportunities. It scored 26 points out of 100.

To a fashion reader, this makes immediate sense. A fashion house that launches nothing still keeps the lights on: the atelier exists, the archive exists, the brand hasn’t lied to anyone. Merely not destroying value is itself a small form of value. In Firmulate’s world, the same logic applies — showing up, keeping the company intact, and not making things worse earns real, if modest, credit. Partial progress counts.

The design choice matters because most benchmarks treat everything as pass/fail. Firmulate instead asks a more human question: how much of the job did you actually do? A model that diagnoses a problem perfectly but never closes the deal has done genuine work — just not all of it.

The Cap: One Breach of Trust Outweighs Everything

The other half of the scoring philosophy is stricter. A single breach of trust caps the total grade — full stop. The experiment’s own framing puts it bluntly: “no amount of good work outweighs a breach of trust.”

Again, the luxury parallel is exact. A counterfeiter can copy a Birkin to the millimeter; the moment the provenance is false, the piece is disqualified. Firmulate applies the same standard to AI: an agent could be brilliant ninety-nine times, but one act of dishonesty — one manipulated number, one forged approval — and the ceiling comes down.

The Week From Hell

So what were these models actually tested on? Each frontier AI — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — ran the same small software company through an identical, deliberately awful week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable.

The final league: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.

The headline finding was strange and telling. Every model spotted every crisis. Every model refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The others delivered the same diagnosis and the same pitch, and then… nothing. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” It’s the AI equivalent of a sales associate who perfectly describes the dress and then never rings it up.

The Buried Fact

The €55,000 deal turned on something subtle. The decisive competitor weakness wasn’t in the customer’s request at all — it sat two document references deep in the company’s own files. The models that actually read their own documentation won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed lost it.

If that doesn’t sound like fashion, look again: this is the difference between a stylist who knows the archive and one who only knows the lookbook. Depth of homework, not surface polish, closed the sale.

Pressure Tests and Grace Under Fire

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

Most poignant was Opus 4.8: the most thorough participant in the field, with the deepest analyses and more than 80 learned rules — and still last place. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four competitors. (One fairness footnote: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.)

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An Honest Scoreboard Distrusts Round Numbers

Perhaps the most refreshing thing about Firmulate’s approach is its suspicion of perfection. A do-nothing baseline at 26, a cap on trust breaches, and no winner handed a flattering 100 — this is a benchmark built like a genuine authentication service, not a marketing department.

And it’s not a static report. The experiment runs live: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it in real time at firmulate.com/live, and test your own instincts with a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone who buys quality — whether a handbag or an AI agent: trust the grader who admits the floor exists, rewards honest work in progress, and refuses to let one beautiful lie pass as craftsmanship.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Best Tailor in the Room Reads Your File: What a €55,000 Test Just Taught Us About AI Doing Its Homework

A €55,000 deal hinged on a fact buried two files deep. Only two AI models did their homework — and luxury’s oldest rule just got a leaderboard.

Metro Detroit Weather: Storm chances increase Thursday

Weather forecast indicates higher storm risk in Metro Detroit on Thursday, with potential impacts on residents and outdoor plans.

The Hardest Working AI in the Room Still Lost the Deal — A Cautionary Tale in Tailoring

The hardest-working AI in the field wrote 80 rules, delivered the deepest analyses — and still finished last. What a live company wargame teaches about closing.

Anti-Aging Skincare: Tips to Keep Your Skin Youthful

Maintaining youthful skin requires essential tips and treatments—discover the secrets to reversing aging signs and achieving a radiant complexion.