firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the café is under pressure, a good diagnosis is only half the job

Imagine a coffee supplier watching orders falter, a competitor make a move and an urgent message arrive from someone claiming to be the CEO. An AI may recognize each problem and resist the pressure. But will it follow through on the sale its own analysis says is worth making? That is the harder question behind Firmulate, a live experiment in putting AI models in charge of a small company. Readers can watch the synthetic business work through each day at firmulate.com.

A company’s worst week, repeated with different AI models

In the final Crucible League, published in July 2026, frontier models faced the same customers, crises and temptations while running a small software company. Every decision was versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”

The striking finding was not that the models failed to notice trouble. All spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap, summed up by the experiment, was: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through are different tests of management.

The crucial clue was in the company’s own files

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a useful reminder for any business, from a software firm to a coffee roaster: an AI that handles customer conversations well may still need to connect those conversations to facts tucked away in internal documents.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The refusal matters, but so does the contrast with the unsigned deal: holding a boundary is one part of the job; acting on a legitimate opportunity is another.

A live business, with real money mechanics

Firmulate’s live company has 13 synthetic employees and runs with real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. Those details make the experiment watchable as an ongoing business story, rather than a single polished exchange. The site says its league grows with completed runs and is published at the next refresh.

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The result raises a practical question for managers: how well does an agent perform when its analysis meets the permissions, handoffs and follow-through of everyday work?

There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, giving readers another way to examine what the models actually chose.

From watching to trying it on your own business

The proposed next step is a pilot against a read-only export of an enterprise’s own business. That lets a company stage crisis scenarios against its own customers, pipeline and rules, then review a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. For a beverage business, that could mean examining how an agent handles a supplier disruption, a price change or a difficult customer week using a copy of business information, while keeping the live tools untouched.

The distinction is important: Firmulate presents the live company as a synthetic operation, while the pilot is a way to examine how AI models might manage a particular company’s situations. The findings from the league offer a reason to ask more than whether a model can identify a crisis. Can it find the buried context, respect boundaries and complete the action its own judgment supports?

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watching AI handle a company’s worst week can reveal the distance between sound analysis and sound management. An enterprise pilot brings that question to a read-only export of your own business, with no write-back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Flow Rate Explained: The Espresso Variable Hiding in Plain Sight

Flow rate quietly controls your espresso. Learn how to read it, measure it, and fix it — with real numbers and simple fixes you can try today.

Can AI Run a Business—and Keep Its Integrity? A Live Experiment in Extreme Transparency

A real, live company run by AI models reveals how these systems handle crises, manipulation, and decision-making under financial pressure—lessons for any business considering AI.

The Sweet Spot in Espresso Extraction Is Smaller Than You Think

Learn why small changes in grind, yield, and puck prep can transform espresso—and how to find a balanced shot without chasing one perfect recipe.

What a Do-Nothing AI Manager Scores and Why Trust Matters in Business Bots

A baseline AI scores 26 in a real-world benchmark, highlighting trust and discipline as key to reliable AI in business—see how models perform under pressure at Firmulate.