
Imagine deploying an AI assistant in your cafe that not only handles orders but also manages crises, reads critical files, and makes trustworthy decisions. Surprisingly, even the most indifferent AI gets a baseline score of 26 points in a rigorous industry benchmark, highlighting the importance of trust and accountability in AI performance.
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Importance of Real-World AI Testing
While many AI demos focus on how well they chat or generate content, the true test lies in how they handle complex, high-stakes scenarios—like managing a business during its worst week. The recent Firmulate Crucible League experiment subjected four frontier AI models to a simulated software company crisis, with real money mechanics, customer crises, and manipulation attempts. This approach offers a clear picture of how AI behaves under pressure, beyond superficial chat quality.
As an affiliate, we earn on qualifying purchases.
Why a Do-Nothing Baseline Scores 26
In this benchmark, even a model that does nothing scores 26 points. This isn’t a flaw but a reflection of the scoring methodology: partial progress counts, and trustworthiness caps the total grade. For instance, a model that refuses manipulation attempts and reads critical files can close real deals, even if it misses some opportunities or slips in process discipline. The score of 26 indicates a minimum level of honesty and diligence embedded in the system, establishing a floor for meaningful AI performance.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed
- Crises and Manipulation: All four models identified every crisis and refused manipulation attempts, such as fake CEO messages and reporter tricks. Kimi K3 justified its refusals with a reason that treats suspicious requests as impersonation risks.
- The Hidden Weakness: The decisive factor was reading in-depth company files. Reading two document references deep in the company’s own files led to closing a €55,000 deal — the best outcome. Models that accessed this information succeeded; those that did not missed the opportunity.
- Discipline and Process: Opus 4.8, the most thorough participant, left the close on the table and showed discipline slips during execution. It failed to escalate issues into the proper departments, reducing its final score despite its analytical depth.
- Honesty Under Pressure: The models’ ability to resist manipulation and read vital internal information was paramount. The only breach that capped the score was a single trust violation; no model exceeded this limit, emphasizing that trustworthiness is non-negotiable in business settings.
AI trustworthiness testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Business Decision-Makers Should Care
For owners and managers of cafes, bakeries, or beverage shops considering AI assistants, the lesson is clear: it’s not just about how well an AI writes or responds. It’s about whether it can finish what it starts, read your critical files, and stay honest under pressure. An AI that can’t be trusted to avoid manipulation or read key documents risks making costly mistakes or missing golden opportunities.
As an affiliate, we earn on qualifying purchases.
The League Table: How Models Compare
- gpt-5.6-sol: scored 95, found the buried fact, closed the deal—full performance.
- Kimi K3: scored 93, closed the deal with the cleanest discipline.
- Sonnet 5: scored 88, closed the deal but with slightly more process slips.
- Fable 5: scored 77, also closing the deal but with more slips and missed discipline.
These results show that even with similar outputs, trustworthiness and discipline are what set the top models apart. The significance isn’t just in the score but in the model’s ability to stay honest and thorough when it counts most.
Practice with Firmulate
Business leaders can run their own AI ‘wargames’ against a read-only export of their enterprise data—ensuring the AI can handle crises, read files, and resist manipulation before deploying it in the wild. This testing isn’t hypothetical; it’s live, accessible, and transparent at Firmulate.

Trustworthiness, thoroughness, and discipline are critical for AI to add real value in business. The Firmulate benchmark reveals that even a do-nothing baseline scores 26 points, setting a meaningful floor for honest AI performance essential for any customer-facing or decision-making role.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
