
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the café is under pressure, a good diagnosis is only half the job
Imagine a coffee supplier watching orders falter, a competitor make a move and an urgent message arrive from someone claiming to be the CEO. An AI may recognize each problem and resist the pressure. But will it follow through on the sale its own analysis says is worth making? That is the harder question behind Firmulate, a live experiment in putting AI models in charge of a small company. Readers can watch the synthetic business work through each day at firmulate.com.
A company’s worst week, repeated with different AI models
In the final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations while running a small software company. Every decision was versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
The striking finding was not that the models failed to notice trouble. All spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap, summed up by the experiment, was: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through are different tests of management.
The crucial clue was in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a useful reminder for any business, from a software firm to a coffee roaster: an AI that handles customer conversations well may still need to connect those conversations to facts tucked away in internal documents.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The refusal matters, but so does the contrast with the unsigned deal: holding a boundary is one part of the job; acting on a legitimate opportunity is another.
A live business, with real money mechanics
Firmulate’s live company has 13 synthetic employees and runs with real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. Those details make the experiment watchable as an ongoing business story, rather than a single polished exchange. The site says its league grows with completed runs and is published at the next refresh.
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The result raises a practical question for managers: how well does an agent perform when its analysis meets the permissions, handoffs and follow-through of everyday work?
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, giving readers another way to examine what the models actually chose.
From watching to trying it on your own business
The proposed next step is a pilot against a read-only export of an enterprise’s own business. That lets a company stage crisis scenarios against its own customers, pipeline and rules, then review a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. For a beverage business, that could mean examining how an agent handles a supplier disruption, a price change or a difficult customer week using a copy of business information, while keeping the live tools untouched.
The distinction is important: Firmulate presents the live company as a synthetic operation, while the pilot is a way to examine how AI models might manage a particular company’s situations. The findings from the league offer a reason to ask more than whether a model can identify a crisis. Can it find the buried context, respect boundaries and complete the action its own judgment supports?

Put your own playbooks to the test
Watching AI handle a company’s worst week can reveal the distance between sound analysis and sound management. An enterprise pilot brings that question to a read-only export of your own business, with no write-back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
