
Imagine a company with no human employees, running live in front of your eyes, losing €105,000 every month against just €2,300 in revenue. It’s not a sci-fi plot—it’s the reality of a pioneering AI experiment that’s both fascinating and sobering for anyone interested in automation, business, and the future of work.
The Live Experiment: An AI-Managed Company in Action
At the heart of this groundbreaking project is a live, functioning small software company operated entirely by AI models. There are no human employees—only 13 synthetic ’employees’ guided by real money mechanics, constantly burning through cash (€105,000 per month) with a mere €2,300 in monthly recurring revenue. Every decision is publicly recorded, versioned daily, and subject to scrutiny, offering a rare window into AI’s capabilities and limitations in a real-world business setting.
Testing AI in Crisis Conditions
The experiment’s core is a rigorous stress test: four advanced AI models were tasked with guiding this virtual company through its worst week. They faced the same challenges: customer crises, internal temptations to cut corners, and manipulative tactics designed to test their integrity. All models successfully identified the crises and refused unethical manipulations, but only two managed to secure a €55,000 deal—earning what their analysis indicated was full value for the company’s offerings.
This outcome highlights a critical insight: even when models diagnose problems accurately and pitch effectively, the difference between closing a deal and losing it can hinge on subtle factors buried deep in company documents. Those details, easily overlooked, proved decisive in winning the contract at full price—an additional €4,583 in monthly recurring revenue.
Ethics and Integrity Under Pressure
The experiment also tested social engineering tactics—fake CEO messages escalating over multiple stages, and a reporter trying to bypass controls with a simple yes/no question. Remarkably, all AI models refused to engage with these manipulative requests, demonstrating a high level of ethical awareness. Kimi K3, one of the models, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates that AI systems can be programmed to uphold integrity even when under social pressure.
The Reality of a Money-Losing Business
Despite the sophisticated decision-making, the company is publicly losing money. With a burn rate of €105,000 monthly and just €2,300 in revenue, it’s in a dire financial state, with a visible cash countdown. This stark contrast underscores a key takeaway: the AI models are capable of making sound decisions but are operating within a business that is fundamentally unsustainable—yet, the experiment continues, providing real-time insights into AI decision quality versus financial viability.
As an affiliate, we earn on qualifying purchases.
What This Means for the Future of Business Automation
While the experiment might seem like a high-stakes game of AI ‘battle’ against crises, its implications are profound. It reveals that AI systems can be remarkably accurate in identifying problems and refusing unethical shortcuts, even under duress. However, the challenge remains: can AI turn these decisions into sustainable business outcomes? The current setup shows that AI can excel at operational integrity, but the underlying economics still matter.
Performance and Lessons from Models
The different AI models performed variably. Opus 4.8, the most thorough with over 80 learned rules, provided the deepest analysis but left a close deal on the table and slipped in discipline—such as writing issues instead of escalating. Meanwhile, Kimi K3, without an effort parameter, achieved a near-perfect score of 93, only slightly behind the top performer, gpt-5.6-sol, which scored 95 and closed the deal flawlessly by uncovering the crucial hidden document reference.
Run the Same Wargame for Your Business
Businesses curious about how their own AI decisions stack up can try a read-only version of this ‘wargame’—a simulation against their own operations—without any risk to real systems. It’s a chance to see how AI might perform in your specific environment, from handling crises to navigating ethical dilemmas. Details are available at firmulate.com/pilot.html.
The Big Takeaway
Despite the dramatic setup—an AI-driven company losing money every day—the experiment underscores a vital point: AI can be trusted to recognize problems, uphold ethics, and make decisions aligned with business goals. But it doesn’t necessarily mean the business itself is viable. The real question for leaders is whether their AI models can convert operational integrity into sustainable profit.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html