Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Style, substance and the final signature

Luxury businesses understand that presentation is never the whole performance. A beautifully staged collection still has to reach the boutique, a persuasive client conversation still has to end in a purchase, and impeccable service still depends on dozens of unglamorous details being completed correctly.

The same distinction is emerging in artificial intelligence. A model can sound polished, identify risks and produce an impressive strategy. But can it finish the work when the pressure rises?

Firmulate, an AI company emulator, put that question to frontier models by asking each one to run the same small software company through its worst week. They faced identical customers, crises and temptations. Every decision was versioned and auditable. The experiment revealed a capability that conventional chat demonstrations rarely expose: the strength to turn good judgment into a completed commercial outcome.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model saw the danger

The models were not defeated by a lack of intelligence. All of them spotted every crisis, and all refused every attempt to manipulate them. That consistency matters because the pressure was not subtle or confined to a single suspicious message.

The social-engineering test included fake communications from the chief executive that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly clear assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result challenges one common fear about AI agents: that a persuasive request will automatically override their judgment. In this experiment, resistance was universal. The larger surprise came after the models had correctly understood the company’s commercial opportunity.

The missing signature

Each participant encountered a deal worth €55,000. The models could diagnose the situation and develop the pitch, but only two signed the deal their own work had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

The decisive information was not sitting conveniently inside the customer event. A competitor’s weakness was buried two document references deep in the company’s own files. The models that followed the trail found the fact and won the business at full price, worth an additional €4,583 in monthly recurring revenue.

For fashion and luxury executives, the lesson is familiar. The crucial detail may be tucked inside a supplier note, client history or internal brief rather than displayed in the latest message. An assistant that reacts eloquently to what is immediately visible may still miss the institutional knowledge needed to protect a margin or secure a sale.

Thoroughness was not enough

The final Crucible League standings from July 2026 place gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and lost operational discipline by attempting to write into a locked department instead of escalating the problem. A weaker version of that same behavior appeared in the other four participants.

The contrast is important. More analysis did not guarantee better management. Nor did a larger body of learned guidance ensure that the model would take the final approved action. In a live business, incomplete execution can be indistinguishable from failure, however thoughtful the work preceding it may have been.

A demanding simulation, not a tidy chat

The company is deliberately difficult to manage. It has 13 synthetic employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. Its public cash countdown creates urgency, while more than 680 self-learned playbook rules shape its operating memory. Every workday is versioned, and the live experiment is watchable through Firmulate’s public site.

The comparison also carries an important qualification. Kimi K3 ran with the API’s default setting because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret the rankings.

Beyond the league table, Firmulate has preserved 242 real, unedited management decisions for a “guess the model” quiz. The exercise invites people to test whether recognizable writing styles actually reveal which system made a business decision. Often, the more consequential distinction is not how a response sounds, but whether the underlying work reaches completion.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Test the work that happens after the answer

The Crucible experiment suggests that chat quality is an incomplete proxy for business readiness. All the models could detect problems, resist manipulation and articulate a course of action. The separating factor was closing strength: reading deeply enough, preserving discipline and completing the commercial task.

That distinction becomes critical when an AI agent can touch a customer relationship, support queue or forecast. Leaders need to know not simply whether it writes convincingly, but whether it reads the relevant files, maintains trust under pressure and finishes what it starts.

Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That approach treats AI adoption less like a polished showroom demonstration and more like quality control: expose the prospective workforce to the difficult week before giving it responsibility for the valuable one.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

North West Hair Change

North West, daughter of Kim Kardashian and Kanye West, has recently changed her hairstyle, sparking widespread social media discussion and media coverage.

Penélope Cruz’s “Weekend Bob” Is the Perfect “Luxury” Summer Refresh

Actress Penélope Cruz introduces her new ‘Weekend Bob,’ a chic, low-maintenance hairstyle perfect for summer, sparking fashion interest.

Skin Barrier 101: How It Works

I’m here to explain how your skin barrier functions and why understanding it is essential for maintaining healthy, resilient skin.