
Imagine planning a week-long outdoor adventure—packing the right gear, choosing the best routes, and handling unexpected setbacks. Now, picture AI managing this trip, making crucial decisions on the fly. Would it be reliable? Would it stay honest and stay the course? The same questions apply to AI managing real businesses. A groundbreaking live experiment is now revealing how different AI models behave in high-stakes management scenarios.
The Experiment: Putting AI to the Test in a Live Business Environment
At the heart of this unfolding experiment is a small software company facing its worst week—crises, temptations, and tough decisions that test the integrity and judgment of AI management systems. Four leading frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with running this company through simulated crises, all in real time. Every decision was recorded and made auditable, providing a rare glimpse into each AI’s personality and management style.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Different Approaches, Same Challenges
Despite the models being different in structure and training, they all successfully identified every crisis and refused manipulation attempts. For example, social engineering tests involving fake CEO messages and a reporter trick—the models refused all five manipulation attempts, with Kimi K3 explicitly stating, “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a shared capacity for integrity under pressure.
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Who Sealed the Deal?
While all models detected crises and maintained honesty, only two models actually signed the €55,000 deal that their analysis had earned. The models that read deeper into the company’s records and identified hidden vulnerabilities managed to close the deal at full price, worth an extra €4,583 MRR. This emphasizes that reading and understanding company files can be a decisive factor—a detail often invisible in traditional chat demos.

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Weakness and Discipline Gap
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ultimately finished last. It left the deal on the table and showed lapses in discipline—such as writing attempts into a restricted department instead of escalating issues. Interestingly, all models displayed similar weaknesses, suggesting that depth of analysis alone doesn’t guarantee perfect performance.
As an affiliate, we earn on qualifying purchases.
Understanding the Personalities of AI Models
These models exhibit measurable management personalities. For instance, Kimi K3 ran without an effort parameter, resulting in a more conservative, disciplined approach. Conversely, Opus 4.8’s extensive rules didn’t prevent discipline slips, indicating that personality traits matter more than raw capacity.
Why This Matters for Outdoor Enthusiasts and Business Leaders Alike
Whether you’re planning a trek through rugged terrain or managing a high-stakes project, the core question remains: Will your AI system finish what it starts? Will it stay honest under pressure? And how much useful work does it deliver? This experiment shows that AI behavior is not just about performance scores but about personality—trustworthy, disciplined, or disengaged.
Experience the Live Wargame
For those curious to see these decisions unfold, the live site allows you to watch the company in action—battling crises, making decisions, and revealing the personalities behind the models. You can even run your own simulations against your business data, all without risking real systems. Visit firmulate.com/quiz.html to explore how AI models perform in scenarios that matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html