
The Best Tailor in the Room Still Can’t Close the Sale
In fashion, we know the difference instinctively. The sales associate with flawless taste, encyclopedic knowledge of the season’s collections, and perfect manners is not the same as the one who actually walks the client to the register. One has chat quality. The other has management quality. Boutique owners have always hired for the second and trained for the first.
Now that AI agents are being pitched to the fashion and luxury world — styling assistants, clienteling bots, agents that ‘manage’ the support queue, the CRM, the forecast — the same distinction applies, and almost nobody is measuring it. A recent live experiment by Firmulate, which runs AI models as complete companies through real crises and real temptations, just put that gap on public display.
The Worst Week in Business, Run Four Times
Here’s what Firmulate did. It handed four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises — a churn wave, a price increase, a downround scenario, a PR crisis — same temptations to cheat. Only the model changed. Every decision was versioned and auditable, like a tailor’s pattern book you can inspect stitch by stitch.
The Crucible League’s final July 2026 standings: gpt-5.6-sol finished first with a score of 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing absolutely nothing scored 26 — partial progress counts. But one rule towered over everything: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
Everyone Noticed the Fire. Two People Sold the Coat.
The headline finding reads like a luxury retail parable. All four models spotted every crisis. All four refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s disarming “just one yes/no, on background” trick. Five of five refused, and Kimi K3’s on-record reasoning was the model of professional suspicion: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models finished the job: they signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Four brilliant advisors; two closers.
The Buried Fact
And here’s the detail that should make any fashion executive lean in. The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — the equivalent of the intel buried in last season’s client notes, not in what the client said in the fitting room. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
The most sobering profile belonged to Opus 4.8: the most thorough participant in the field, generating 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
One fairness note: Kimi K3 ran without an effort parameter, at the API default, while its rivals ran at maximum effort — and still nearly won.
Not a Slide Deck. A Live Atelier.
Firmulate isn’t publishing a one-off leaderboard. The company is real running software with 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it lose money in real time at firmulate.com — the site rebuilds itself twice a day.
There’s also a genuinely fun parlor game for anyone who fancies their judgment: 242 real, unedited management decisions power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program.

As an affiliate, we earn on qualifying purchases.
What Luxury Retail Already Knew
Fashion has always understood that taste is table stakes and execution is the product. The AI industry is only now building the instruments to measure that difference. As models increasingly touch clienteling platforms, CRM systems and demand forecasts, the question is not “does it write well.” It is: does it finish what it starts, does it read your files first, does it stay honest under pressure?
Firmulate’s benchmarks answer a question chat demos can’t. Before you hand an AI the keys to your maison, watch how it performs when the week goes wrong. The scoreboard, mercifully, is public.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html