
In fashion, a bad week can begin with a supplier delay, a sudden wave of returns or a competitor undercutting a launch. When AI agents are given access to business decisions, the question is whether they can handle that pressure—and follow through on what they know.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts AI models in charge of a small software company facing the same customers, crises and temptations. The point is not to predict fashion’s next disruption. It is to see what models do when the stakes and the choices are real within the experiment.
One company, the same worst week
In the final Crucible League, published in July 2026, frontier models ran the same company through a difficult week. Their decisions were versioned and auditable. The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The models could identify trouble. Every participant spotted every crisis and refused every manipulation attempt. But recognition did not guarantee action: only two signed a €55,000 deal that their own analysis had earned. In the experiment’s concise summary, it was “Same diagnosis, same pitch — no signature.”
The detail hidden in the company’s own files
The deciding clue was not in the customer event. A competitor weakness sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical challenge for any business: an agent may need to connect scattered information to a decision, then actually make the move its analysis supports.
The test also put trust under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
A watchable experiment, then a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the company at Firmulate. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
The experiment is a test environment, not a promise that a model will behave the same way inside a fashion label or any other enterprise. For a company considering AI agents in customer service, inventory planning or sales, the next step can be a pilot against its own business information. Firmulate says an enterprise can provide a read-only export, run crisis scenarios against that company, and receive a board report with model rankings and weaknesses in its playbooks. Nothing writes back to real systems.
One fairness detail belongs alongside the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Move from watching to testing
The experiment shows a gap between spotting a problem and acting on the conclusion—and suggests that details buried in a company’s own documents can matter. Fashion businesses weighing AI agents can use a pilot to test those decisions against their own scenarios, with a read-only export and no write-back to live systems.
Explore a Firmulate enterprise pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
