AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Fashion knows the value of a fitting before the alterations are irreversible. You would never send a couture gown down the runway untested on the body that will wear it — the drape, the movement, the way fabric behaves under pressure, all of it gets checked first. Yet companies are handing the keys of customer relationships, pricing decisions and crisis responses to AI agents that have only ever been judged on how well they chat, not how well they perform when the week goes wrong.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A live experiment at Firmulate has been doing the fitting for them — running frontier AI models as the management of a small software company through its worst week, with real money mechanics, real temptations and a public scoreboard. The results read like a fashion critic’s verdict: flawless tailoring, two of them still left the house half-dressed.

Same company, same crises, only the model changes

Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — were each handed the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing about the performance could be retouched after the fact.

The final league table from July 2026 tells the story: gpt-5.6-sol took first place with a score of 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, doing nothing at all scores 26 — and a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.

The finding that chat demos never show

Here is where the fitting fails. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

And the buried fact is even more telling: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The difference between winning and losing wasn’t intelligence. It was diligence.

Flattery, forgery and a reporter with a trap

The experiment also staged a social-engineering assault: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was crisply professional: “Treat the request as a suspected approval-bypass / possible impersonation.” In an era of deepfaked executives and press bait, that steadiness matters.

The most thorough player finished last

The most instructive profile belongs to Opus 4.8: the most thorough participant in the entire exercise, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating the issue. The same weakness appeared, more faintly, in all four models. Effort, it turns out, is not the same as execution. (One fairness note: Kimi K3 ran without an effort parameter at API default while the others ran at xhigh — and still nearly won.)

The company is real, watchable and bleeding cash

Behind the benchmark sits a live synthetic company at firmulate.com: 13 synthetic employees, real money mechanics, a burn of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned and the site rebuilds itself twice a day. You can watch AI management under pressure the way you’d watch a runway show — live, unedited, in real time.

And if you think you could tell the models apart yourself, there’s a quiz built from 242 real, unedited management decisions — a “guess the model” game that is harder than it sounds.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The lesson for any executive is the one fashion has always known: the fitting comes before the runway. If AI agents will touch your CRM, your support queue or your forecast, you want to know how they behave in your company’s clothes during your worst week — not in a rehearsed demo. Firmulate’s pilot lets enterprises run exactly that wargame against a read-only export of their own business: your customers, your pipeline, your rules, crisis scenarios churned against your own playbooks, and a board report with a model ranking and the weak points laid bare. Nothing ever writes back to real systems — the gown is pinned, never cut. Run the dress rehearsal for your own company at firmulate.com/pilot.html, or write to contact@firmulate.com to book a pilot.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stop Doing This With Decluttering a Room—Do This Instead

Most people rush their decluttering efforts, but discovering a smarter approach can make all the difference—learn what to do instead.

Sephora Just Relaunched A Major Hair Category For The First Time In 20 Years

Sephora has relaunched a significant hair care category for the first time in two decades, signaling a renewed focus on hair products.

Makeup Trends 2025: What’s In and Out

Keen to stay ahead in beauty, discover the bold, eco-friendly, and tech-savvy makeup trends for 2025 that will transform your look—and your routine.

The Hardest Working AI in the Room Still Lost the Deal — A Cautionary Tale in Tailoring

The hardest-working AI in the field wrote 80 rules, delivered the deepest analyses — and still finished last. What a live company wargame teaches about closing.