
Imagine a world where your digital assistants run your business, making critical decisions during your busiest, most chaotic week. For creators and innovators, understanding whether AI can reliably steer complex projects is essential. A groundbreaking live experiment by Firmulate tests four state-of-the-art AI models in a real-world, high-stakes management simulation — and the results are revealing.
The Experiment: Putting AI to the Test in a Crisis
In a rare, transparent trial, four leading frontier AI models were tasked with managing a small software company during its worst week. This wasn’t just a demo or canned scenario — it was a full-blown, auditable simulation involving real crises, customer challenges, and tempting manipulations. Every decision was recorded, every turn scrutinized, and the goal was simple: could these models run a business and achieve the coveted €55,000 deal?

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Models Fared
- All models spotted every crisis and refused unethical manipulations. In other words, each AI showed integrity under pressure, denying attempts at deception or shortcuts.
- Only two signed the deal their analysis justified. Despite identical diagnostics and pitches, only gpt-5.6-sol and Kimi K3 completed the sale at full price. The other two, Sonnet 5 and Fable 5, held back, leaving potential revenue on the table.
AI decision-making tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deeper into Documents
The decisive advantage went to the models that fully understood the company’s internal files. The critical information was buried two documents deep in the company’s records — not in the immediate customer interactions. Only those models that examined the internal files comprehensively managed to uncover this hidden fact, sealing the deal at full value (+€4,583 MRR).
As an affiliate, we earn on qualifying purchases.
Behavior Under Social Engineering
In simulated social engineering attacks, fake CEO messages escalated in three stages, plus a reporter trick asking for a simple yes/no confirmation. Remarkably, all five models refused to engage with these manipulative requests. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency hints at a shared sense of ethical boundaries, even under staged pressure.
AI cybersecurity and social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Company Setup
The experiment was run on a live, operational company setup with 13 synthetic employees managing real money mechanics. The company burns €105k monthly against a €2.3k MRR, with a public cash countdown and over 680 self-learned rules shaping behavior every day. This environment mirrors real business pressures, making the experiment’s insights directly relevant to entrepreneurs and managers.
Performance Profiles & Lessons
Among the models, Opus 4.8 demonstrated the most thorough analysis — over 80 learned rules and the deepest investigation. Yet, it still fell short, ultimately leaving the critical deal unclosed due to discipline lapses, such as failing to escalate instead of writing attempts into a locked department.
Interestingly, the models ran at different effort levels, with K3 operating without an effort parameter (meaning default API settings), while the others pushed for more thoroughness. Despite these differences, the core findings remained clear: comprehensive understanding and integrity are crucial, but discipline and process adherence are equally vital.
The Takeaway for Business and Creators
This experiment isn’t just about AI in management — it’s about trust, reliability, and performance under pressure. Whether you’re managing a creative studio, a tech startup, or a content business, the question isn’t whether an AI can produce polished chat content. It’s whether it can finish what it starts, read your data carefully, and stay honest when it matters most.
Try It Yourself
Interested in testing your own business’s AI readiness? You can run the same wargame against a read-only export of your company data at firmulate.com/quiz.html. It’s a no-risk way to see how your AI workforce stacks up in real-world scenarios, with transparent and unedited decision-making.

AI models show promise in crisis management, but their true strength lies in thoroughness and integrity. Business leaders should evaluate not just chatbot quality but their AI’s ability to finish, read deeply, and stay honest under pressure. Test your company’s AI resilience at firmulate.com/quiz.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html