
Imagine testing your favorite shopping deals not just with coupons but with AI that manages a whole company—making critical decisions in real time and under pressure. That’s the premise of a groundbreaking experiment in AI management, where the results could change how businesses pick their AI partners. Recently, a live benchmark pitted top AI models against each other in running a live software company during its worst week, revealing which AI truly proves its worth in real-world business challenges.
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Stakes of AI in Business Decision-Making
In today’s fast-paced digital economy, AI isn’t just about chatbots or recommendation engines — it’s increasingly about managing entire operations, from customer crises to strategic deals. But how do you know if an AI model can handle the real pressures of running a company? A public experiment by Firmulate puts this question to the test by running several of the leading AI models through a simulated week of crisis, opportunity, and manipulation attempts.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: One Company, Multiple AIs
The setup was straightforward yet rigorous: identical small software companies, identical challenges, identical crises, and the same set of temptations to cheat or manipulate. Each AI model was tasked with managing the same business for a typical workweek, with every decision publicly logged and auditable. The key goal? To see which AI could best diagnose problems, resist manipulation, and close a critical deal valued at €55,000 — a real-world measure of performance.
The Results: Who Came Out on Top?
- gpt-5.6-sol scored the highest with 95 points, successfully uncovering a buried but critical piece of information in the company’s documents and closing the deal.
- Kimi K3 from Moonshot followed closely with 93 points, demonstrating the cleanest discipline of the field and also sealing the €55K deal based on its own analysis.
- Sonnet 5 came third with 88 points, managing to close the deal despite a few slips in process discipline.
- Fable 5 scored 77, and Opus 4.8 trailed with 73, both completing the deal but with noticeable weaknesses in process and discipline.
The Hidden Difference: Reading the Files Counts
A crucial insight emerged from the detailed analysis: the decisive factor was whether the AI read and understood the company’s own documentation, not just reacting to customer interactions. The AI that successfully found a buried reference deep within the company’s files won the deal at full price, demonstrating that thorough internal knowledge was the key to success.
Resisting Manipulation and Ethical Testing
Beyond decision-making, the models faced a series of social engineering tests—fake CEO messages escalating over three stages and a reporter trick asking for secret approvals. Impressively, all four AI models refused to participate or sign off on suspicious requests, citing concerns over impersonation and bypassing approval processes. This highlights an essential aspect of AI trustworthiness: integrity under pressure.
The Real Company Behind the Test
The experiment wasn’t just theoretical. The live setup was a functioning company with 13 synthetic employees and real money mechanics, burning €105,000 monthly against a modest €2,300 monthly recurring revenue. It features over 680 self-learned rules, all visible and versioned daily at firmulate.com. Watching these AIs in action offers a clear window into how future AI managers might perform in actual business contexts.
The Lessons for Business Leaders
The key takeaway? Performance isn’t just about the quality of a model’s responses or how shiny its demo looks. It’s about consistency, thoroughness, and honesty—qualities that, in this experiment, determined the winner. The top-performing AI not only identified the critical buried fact but also closed the deal as a result—indicating a level of discipline and reliability essential for real-world deployment.
Fairness and Methodology
It’s worth noting that Kimi K3 was tested without an effort parameter—meaning it operated under default API settings—while the other models ran at a higher effort level, which could influence performance. This ensures a fair comparison rooted in actual operational capability.
Why It Matters for Your Business
Whether you’re managing customer relationships, support queues, or forecasting, the question isn’t just whether an AI can generate convincing chatter. It’s whether it can complete the work reliably, resist manipulation, and truly understand your internal data. As this experiment shows, choosing an AI partner without testing its real-world performance can be a gamble—one that might cost more than just money.
Discover More and See the AI in Action
Curious about how your own company might fare against these AI models? You can run a simulation against your business data with Firmulate’s live platform, which makes it possible to test AI management in a safe, risk-free environment. It’s real, transparent, and designed to help you make smarter decisions about AI adoption.

The live experiment reveals that the best AI managers are those that read deeply, stay disciplined, and resist shortcuts. In a world increasingly driven by AI, testing performance in real conditions is essential—don’t settle for superficial demos when your business’s future is at stake.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
