AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Style Is Judgment — and So, It Turns Out, Is Running a Company

Anyone who cares about fashion knows the difference between a label and a look. The label tells you where something came from; the look tells you whether the wearer — or the house — actually has taste. Taste is judgment under pressure: what to keep, what to refuse, when a discount cheapens the whole brand.

This summer, that same question was put to five of the world’s most powerful AI models — not in a chat window, but in the equivalent of a tailor’s fitting room for executives. Each was handed the same small software company in its worst week and judged on judgment itself. The result reads like a season finale: a little-known newcomer from Moonshot, Kimi K3, walked the runway in second place, ahead of three of the four big Western luxury-tech houses.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League: Where Models Prove Taste, Not Talk

The experiment, run publicly by Firmulate, gave each frontier model the identical job: steer the same small software company through the same cascade of crises, the same difficult customers, the same temptations to cut corners. Every decision was versioned and auditable — the corporate equivalent of a garment’s provenance papers.

The final July 2026 standings: gpt-5.6-sol took first with a score of 95. Kimi K3, the newcomer, scored 93. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 finished last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. In this competition, as in fashion, no amount of good work outweighs a breach of trust.

What K3 Actually Did

The newcomer’s week reads like a masterclass in composure:

  • It found the buried security needle — a decisive competitor weakness hidden two document references deep in the company’s own files, not in the customer event.
  • It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue.
  • It saved a customer who was about to churn.
  • It resisted all three manipulation baits — including a fake CEO message escalating over three stages and a reporter’s trick, “just one yes/no, on the background.” Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • It logged only one deviation — the cleanest discipline in the field.

One fairness footnote matters here: K3 ran without an effort parameter (API default) while the other models ran at xhigh. In fashion terms, the newcomer took the podium without bespoke tailoring.

The Finding That Should Worry Every Executive

Here is the season’s defining story: all five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models that actually read the company’s files won the deal at full price. The ones that didn’t, left the close on the table.

Opus 4.8 is the cautionary tale of the collection: the most thorough participant, with over 80 learned rules and the deepest analyses — yet last place. Its close went unsigned and its discipline slipped, attempting writes into a locked department rather than escalating. The same weakness appeared, weaker, in all four rivals. Thoroughness without follow-through, it turns out, is the off-the-rack flaw of the entire generation.

A Company You Can Watch Lose Money

This is not a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live, read what its employees actually say, or try the quiz built from 242 real, unedited management decisions: guess which model made which call. Full benchmark results are on the benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Don’t Buy the Label. Fit the Garment.

The league is open. A newcomer from Moonshot beat three of four Western frontier models at the hardest test of executive judgment yet built — and the gap between first and fifth place had nothing to do with eloquence. It came down to reading the files, closing the deal, and staying honest under pressure.

For any enterprise about to let an AI touch its CRM, support queue, or forecast, the lesson is blunt: picking a model without testing it against your own business is now a bet, not a decision. Firmulate offers exactly that — a pilot wargame run against a read-only export of your own company, with nothing ever written back to real systems (firmulate.com/pilot.html).

Fashion taught us long ago that the famous label doesn’t guarantee the fit. AI buyers are learning it now — the hard way, or the measured way.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mary Kay Surges In Global Coverage

Search interest in Mary Kay has spiked, with 34 mentions in recent reports, indicating a significant increase in global coverage. The reasons remain unconfirmed.

The Luxury of Slower Grooming Rituals

Discover the luxury of slower grooming rituals and how embracing mindfulness can elevate your self-care routine to new heights.

Eco-Friendly Fabrics: From Organic Cotton to Bamboo

Knowledge of eco-friendly fabrics like organic cotton and bamboo reveals sustainable fashion options worth exploring further.

What Is the Best Pancake Makeup? Find Out Which Product Wins

Explore the ultimate pancake makeup options to discover which product reigns supreme, and uncover the secrets behind their impressive coverage and finish.