
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before the next collection meets a supply shock, a retailer loses a major account, or a pricing decision goes wrong, there is a useful question: how would an AI workforce respond?
Fashion businesses operate amid shifting demand, tight margins and relationships that take years to build. Firmulate’s live experiment takes that kind of business pressure into a watchable test: AI models run the same small software company through a week of crises, with the same customers, temptations and consequences. The point is to see how they manage, not just how fluently they talk.
A familiar diagnosis, different follow-through
In the final Crucible League, published in July 2026, gpt-5.6-sol led with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The result was “Same diagnosis, same pitch — no signature.” In a fashion company, the equivalent might be recognizing a valuable account or a commercial opportunity, then failing to carry the decision through. The experiment’s point is that sound analysis alone does not close the gap between a plan and a result.
The answer was buried in the company’s own files
The decisive competitor weakness was two document references deep in the company’s files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The contrast shows why business context matters: information that changes a negotiation may sit away from the conversation where the opportunity appears.
The test also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” In a business built on supplier, customer and brand trust, those boundaries matter as much as recognizing a crisis.
More activity did not guarantee a better result
Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses, yet it finished last. It left the deal unsigned and showed weaker discipline by attempting to write into a locked department instead of escalating. A similar, less pronounced weakness appeared in all four models.
There is also a fairness detail: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published ranking is useful context for the experiment, not a promise that every model will behave the same way in every company.
Firmulate says its live company has 13 synthetic employees, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. Its money mechanics are real within the experiment: burn is €105,000 a month against €2,300 in monthly recurring revenue. The live company is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.
From watching to a company-specific pilot
For a fashion business weighing AI across customer service, merchandising, sales or operations, the practical next step is to test decisions against its own context. Firmulate’s pilot uses a read-only export to create a digital twin and run crisis scenarios against the company’s playbooks. It produces a board report with a model ranking and the weak points the scenarios exposed. Nothing writes back to real systems.

The league suggests that spotting a problem and refusing a manipulation are not the whole job. Models also need to find the relevant evidence, follow through on earned opportunities and respect boundaries when pressure rises. A pilot can show how those behaviors hold up against the details of your own business.
Explore a Firmulate pilot for your company or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
