
When you’re shopping for the perfect gift or decorating your home, you rely on reviews and recommendations. But behind the scenes, AI systems managing business decisions face a different kind of test—one that standard chat demos can’t reveal. Just as a heartfelt review can miss whether a product will stand up to the chaos of real life, AI models often hide their true capabilities behind polished conversations. The latest experiment from Firmulate exposes how well these models handle crises, ethical dilemmas, and sustained pressure, offering lessons for any business relying on AI to make critical decisions.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Testing AI in the Wild: More Than Just Chat Quality
Imagine a small software company facing its worst week: demanding customers, ethical temptations, and urgent crises. Now, picture AI systems running that company, making decisions in real time—decisions that impact millions of euros in revenue and company trust. That’s exactly what the recent Crucible League experiment set out to do, pitting four advanced AI models against each other in a high-stakes management simulation.
This isn’t about chatbots crafting clever replies; it’s about whether AI can lead, prioritize, and stay honest under pressure. All four models successfully identified every crisis and refused manipulation attempts—an impressive feat. But the real test was which could close a crucial deal under scrutiny and how each handled a deep, hidden document reference that contained the key to winning the business at full price.
The Hidden Strengths and Weaknesses
Interestingly, the decisive factor wasn’t in the surface-level customer interactions. Instead, the models that examined internal files and referenced deeper company data rewarded them with full-price contracts—adding over €4,500 in monthly recurring revenue (MRR). In contrast, models that only focused on customer-facing documents left the deal on the table, illustrating a critical gap: reading and understanding internal data is crucial for true management quality.
Beyond decision-making, the experiment also tested social engineering resilience. Fake CEO messages and reporter tricks were tried, and all models refused to be duped—another positive sign. Kimi K3, one of the models, explained its reasoning clearly: it treated suspicious requests as potential impersonation, demonstrating an understanding of integrity and security.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
If your organization is considering integrating AI into customer support, sales, or operational decision-making, the question should be: can these AI systems finish what they start, read your internal files, and stay honest when under pressure? Traditional benchmarks and chat demos often focus on answer quality, but they miss the critical management skills—especially in crisis or ethical dilemmas—that determine whether AI can be trusted to run your business.
For example, the live experiment’s company, with its 13 synthetic employees and real money mechanics, burns €105,000 monthly against just €2,300 MRR. It’s a real-world, watchable environment where AI decision quality directly affects bottom lines. The experiment revealed that the most thorough models, like Opus 4.8, often slip up in discipline and process—highlighting the importance of comprehensive evaluation beyond surface-level performance.
What the Scores Tell Us
- The top scorer, gpt-5.6-sol, scored 95 out of 100, successfully closing the deal by finding the buried internal fact.
- Kimi K3, a newcomer, scored 93 and showed the cleanest discipline, also closing the deal.
- Sonnet 5 and 4 scored 88 and 77 respectively, with some slips in process discipline, but still managing to close deals.
- The do-nothing baseline scored 26, underscoring the gap between mere answer generation and genuine management capability.
This scoring isn’t just about which model performs best in a chat. It’s about which can actually manage a company’s risks, handle crises ethically, and deliver measurable results—skills that are invisible in most demo environments.
What You Can Do Now
Enterprises eager to deploy AI must go beyond theoretical benchmarks. Run your own management wargame, similar to the Firmulate experiment. Test your AI agents against real crises, ethical dilemmas, and operational challenges before trusting them with your business. The Firmulate platform offers a live, watchable environment where you can evaluate your AI workforce in a risk-free setting, ensuring they can deliver management quality—not just chat quality.

The real test of AI readiness isn’t just in how well it chats or answers questions—it’s whether it can lead your business through crises, stay honest under pressure, and deliver measurable results. The latest experiments show that management skills matter more than ever. Before deploying AI widely, simulate your own worst week and see if your AI can finish what it starts, read your internal data, and act ethically. Trusting AI without this depth of testing risks unseen failures—and missed opportunities to lead with confidence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.