firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

The Future of Business Decisions Is AI-Driven — But Which AI Model Does It Best?

Imagine a bustling software company, facing the worst week of its life — tight deadlines, tricky crises, and tempting shortcuts. Now, picture AI models managing these challenges, offering decisions that could make or break the business. Who would you trust? Which AI would stay honest and finish what it starts? Welcome to a groundbreaking experiment where AI models are put through real-world management tests, revealing not just their capabilities but their personalities.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the AI Management Showdown

At the heart of this experiment is Firmulate, a platform that runs AI models as complete companies, complete with real money mechanics, crises, and temptations. Four frontier AI models were tasked with managing a small software company through its toughest week, facing identical scenarios, same customers, and the same pressures to cut corners.

What makes this test fascinating is that all four models identified every crisis and refused every manipulation attempt — a clear sign of their integrity. However, only two managed to close the most critical deal, earning €55,000, by reading the company’s own files and spotting a crucial hidden piece of information. This buried fact, just two document references deep, was the difference-maker—those who read the files won the deal at full price, worth over €4,583 in monthly recurring revenue.

How Models Show Their Personalities

This experiment illustrates that AI models don’t just differ in accuracy; they have measurable management personalities. For example, Opus 4.8, which ran with the deepest analysis and the most rules learned, was thorough but ultimately left money on the table by slipping into unproductive behaviors, like writing attempts into a locked department instead of escalating them. Conversely, Kimi K3, a newcomer with no effort parameter, demonstrated the cleanest discipline and secured the deal.

Another key insight: all models handled social engineering attempts flawlessly, refusing fake CEO messages and reporter tricks — a vital trait for trustworthy AI in real business settings.

Why This Matters to You

If AI agents are going to manage your CRM, support queues, or forecasting, the question isn’t just about how well they chat. It’s whether they can follow through on commitments, read important documents, stay honest under pressure, and deliver consistent value. This live experiment shows that some models, like gpt-5.6-sol, excel in these qualities, while others, despite their analytical depth, may slip into less disciplined behaviors.

For instance, the final league scores reveal the top performers: gpt-5.6-sol scored 95 points, successfully closing the full deal and finding buried information, while Kimi K3 scored 93, also closing the deal with clean discipline. Meanwhile, Sonnet 5 and Fable 5 scored lower, with 88 and 77 respectively, and showed more process slips or left opportunities unexploited.

What You Can Do Now

Firmulate offers a transparent window into this AI management world. You can test your own business scenarios against these models with their interactive quiz that uses real, unedited management decisions. Plus, enterprises can run their own wargames, simulating how AI might handle specific challenges without risking actual business operations (learn more here).

This is not just a demo — it’s a live, watchable experiment happening every business day. The company involved is real, operating with 13 synthetic employees, managing €105K burn each month against €2.3K MRR, and making decisions that matter. You can see the AI’s decisions in real-time, read employees’ actual comments, and even guess which model made which decision.

Infographic —
The findings at a glance — source: firmulate.com.

The Bottom Line: Trustworthy AI Is Possible—and Measurable

This experiment proves that AI models are not just about chat quality—they have personalities, biases, and decision-making styles that can be measured. Some models are disciplined enough to spot hidden risks and close full-price deals, while others might leave money on the table due to process slips. For businesses considering AI management tools, these insights highlight the importance of testing and choosing models that stay honest and finish what they start. Watch this space — and the live company — to see AI’s management personalities in action.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Kalshi promo code CBSSPORTS for Belgium vs. Senegal: Get $15 bonus for 2026 World Cup trading on Wednesday

Use promo code CBSSPORTS on Kalshi to receive a $15 bonus for trading related to the Belgium vs. Senegal match in the 2026 World Cup.

BTGO Investors Have Opportunity To Lead BitGo Holdings, Inc. Securities Lawsuit

BTGO investors have the opportunity to serve as lead plaintiffs in a securities lawsuit against BitGo Holdings, Inc., as detailed in a recent PR Newswire release.

Maple Street Biscuit Company Sold

Maple Street Biscuit Company has been sold to new ownership, ending its previous management. Details on the sale and future plans are still emerging.

ثقة وقوة راسختان في تصنيع FREELANDER 8

شركة تصنيع السيارات تؤكد على ثقتها وقوتها في إنتاج نموذج FREELANDER 8، مع التركيز على الجودة والتكنولوجيا الحديثة.