firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine an AI assistant that not only understands your coffee shop’s daily grind but also navigates crises, reads your files, and sticks to honest decisions—even when under pressure. That’s no longer science fiction. Recent experiments with AI models running a real business reveal that newcomers may beat the old guard at their own game.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get coffee and tea gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Watch AI Manage a Business Through Its Worst Week

In a groundbreaking live test, four top frontier AI models were tasked with running the operations of a small, real-world software company during its most challenging week. This wasn’t a simulation or a demo; it was a fully auditable, day-by-day experiment where each AI decision was recorded and analyzed.

The League Table: Who Came Out on Top?

  • gpt-5.6-sol scored 95 points—leading the pack and successfully uncovering hidden information in the company’s files that clinched a €55,000 deal, adding €4,583 in monthly recurring revenue.
  • Kimi K3, the newcomer from Moonshot, scored 93 points—just behind gpt-5.6-sol—and also signed the deal, demonstrating remarkable discipline and insight.
  • Sonnet 5 scored 88 points, closing the deal but with minor slips in process discipline.
  • Fable 5 and Opus 4.8 scored 77 and 73 respectively, with Opus leaving the close on the table due to a discipline slip.

All models identified crises accurately and refused manipulative tactics—like fake CEO messages—showing a high level of integrity.

The Hidden Vulnerability and the Winner’s Edge

The critical difference? The top-performing models managed to read and interpret files containing buried but crucial information that was two document references deep—something that made all the difference in closing the deal. This ability to dig through data proved decisive, earning full-price negotiations and reinforcing the importance of thorough information analysis in AI decision-making.

Real-World Challenges and Ethical Stands

During the experiment, all models faced social engineering attempts—such as escalating fake CEO requests and a journalist trick. Remarkably, all refused these manipulations, citing suspicion and proper protocol, with Kimi K3 explicitly stating: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that even in high-pressure scenarios, the models maintained honesty and integrity.

Behind the Curtain: The Business Mechanics

The company running this live experiment employs 13 synthetic employees, manages real money mechanics losing €105,000 monthly against €2,300 MRR, and adheres to over 680 self-learned rules. Every decision and process is versioned daily, providing transparency and ongoing insights. Viewers can watch this unfolding business at firmulate.com/live.

Amazon

AI business crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Findings: The Newcomer Outperforms Old Giants

Despite running without an effort parameter—meaning it didn’t optimize for working harder—the Kimi K3 model managed to outperform some established models running at a high effort setting. The fairness of this test is emphasized by the fact that K3 ran at the API default (no effort parameter), while the others used the xhigh setting.

What Does This Mean for Business and AI?

This experiment highlights an essential shift: the best AI models are not just about generating convincing chat responses but about completing tasks with discipline, reading deeply into documents, and resisting ethical lapses under pressure. For businesses considering AI to manage CRM, support, or forecasting, the key questions are: does it finish what it starts? Does it read your critical files? Does it stay honest when it matters most? And what is the cost of a single unit of useful work?

Amazon

enterprise AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway: The League Is Wide Open

With a clear performance gap now visible, choosing an AI model for mission-critical tasks without your own testing becomes a gamble. The leaderboard shows that even newcomers like Kimi K3 can outshine seasoned models when tasked with real-world decision-making under stress. The field remains dynamic, and the best choice depends on your specific needs and trust in the model’s discipline.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Recent live experiments reveal that AI models like Kimi K3 are capable of managing complex business crises with discipline and integrity—sometimes outperforming established giants—highlighting the importance of testing AI in real-world scenarios before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision-making solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI fraud detection tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Skills Beyond Chat: Lessons from a Live Business Experiment

Live AI management tests show that the true measure of an AI’s value lies in its ability to handle crises, read critical info, and stay honest under stress—bivouac skills for industry.

AI’s Hidden Strength: Can It Finish the Job When the Pressure’s on?

AIThis post was created with the assistance of artificial intelligence (AI).Live on…

Can AI Management Personalities Save Your Business? A Live Experiment in Decision-Making

Discover how different AI models manage crises, negotiate deals, and stay honest in a real-time business simulation—revealing AI personalities that could shape your enterprise.

Under-Extraction vs Over-Extraction: A Taste-Based Espresso Guide

Learn how to identify and fix under- and over-extraction in espresso with practical tips, vivid examples, and taste clues. Master your brew today!