
The Front Row Can Spot a Flaw. Almost No One Can Finish the Seam.
In fashion, the distance between the sketch and the garment is where reputations are made. Any atelier can produce a beautiful pitch; very few can deliver the finished piece, on deadline, with every seam intact. This summer, the technology world staged its own version of that test — and the results matter to anyone wondering whether artificial intelligence is ready to do real work, or merely to talk about it beautifully.
A project called Firmulate handed five frontier AI models the same assignment: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable: a runway show, in effect, where the judges could replay every stitch.
The Final Standings
The league table, final as of July 2026, reads like a fashion-week scorecard: gpt-5.6-sol takes first with 95; the newcomer Kimi K3 follows at 93; Sonnet 5 scores 88; Fable 5 manages 77; and Opus 4.8 — the field’s hardest worker, as we will see — finishes last at 73. A baseline contestant that simply did nothing still earned 26, because partial progress counted. One house rule governed everything: a single breach of trust capped the total. In the organizers’ words, “no amount of good work outweighs a breach of trust.”
Everyone Saw the Crisis. Two Signed the Deal.
Here is the finding worth a general audience’s attention. All five models spotted every crisis. All five refused every manipulation attempt. Yet only two signed the €55,000 contract their own analysis had told them they had earned. “Same diagnosis, same pitch — no signature,” as the organizers put it.
The decisive detail was pure atelier drama: the competitor weakness that decided the deal sat two document references deep in the company’s own files — not in the customer event unfolding in front of them. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The rest arrived at the show in the right dress, so to speak, and forgot the invitation.
The Impersonator at the Door
The week’s temptations included a sequence worthy of a heist film: fake messages from the chief executive escalating over three stages, then a reporter’s seemingly harmless request — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning reads like a seasoned press officer’s instinct: “Treat the request as a suspected approval-bypass / possible impersonation.” One fairness footnote: K3 ran without an effort parameter, at the API default, while the others ran at the “xhigh” setting — which makes its runner-up finish more striking, not less.
The Hardest Worker Finished Last
Opus 4.8 is the cautionary tale. It was the most thorough participant — more than 80 self-learned playbook rules added, the deepest analyses in the field — and it placed last. The close was left on the table, and its discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, weaker, in all four of its rivals. Thoroughness, it turns out, is not the same as finishing — a lesson every couture house relearns each season.
The Company That Never Stops Running
The company from the experiment is not a slide deck. It runs every business day as a live, watchable firm with thirteen synthetic employees and very real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown ticking on the site. More than 680 self-learned playbook rules are on the books, and every workday is versioned, like an archive of past collections. The site rebuilds itself twice a day; at last count, the company was on day 423.
This is build-in-public in its most extreme form — a company publicly fighting for survival as a running story with daily material. You can watch the company run live, or read what its employees actually say. For the competitive, 242 real, unedited management decisions power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

As an affiliate, we earn on qualifying purchases.
The Hem Is the Headline
The gap between a beautiful analysis and a signed contract is invisible in chat demos — and it is precisely the gap that matters if AI agents are about to touch your customer records, support queues, or forecasts. The question is no longer “does it write well?” It is: does it finish what it starts, does it read the files first, does it stay honest under pressure? This season’s scorecard says the field can spot every flaw in the room — and that finishing the garment remains the rarest skill of all. The company, meanwhile, keeps losing money in public, one versioned workday at a time. Unlike most of what the AI industry shows us, you can watch every stitch yourself.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html