
Imagine a business where every decision — from managing crises to closing deals — is made by artificial intelligence, all while the company burns through €105,000 every month with only €2,300 in recurring revenue. For creators and tech enthusiasts alike, this isn’t science fiction; it’s the live experiment of Firmulate, a company as real as it is bizarre, and you can watch it unfold every day.
The Live Company That’s Public and Under Pressure
At the core of this experiment is a digital company with 13 synthetic employees, governed by a complex web of 680+ self-learned rules and a public cash countdown. Every workday, the company’s decisions are versioned, auditable, and visible to anyone curious enough to follow along at firmulate.com/live.html. This setup isn’t just a tech demo; it’s a window into how AI models handle real-world business crises and ethical dilemmas under extreme conditions.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI in the Trenches: The Results
Four leading AI models — including the well-known GPT-5.6 and the newcomer Kimi K3 — were tasked with navigating the same brutal week of business challenges. This included managing customer crises, resisting social engineering attempts, and making critical decisions on deals and trust. The results were revealing:
- All four models identified every crisis and refused manipulation attempts, showing a high level of ethical resistance.
- Only two models managed to sign a €55,000 deal their own analysis had earned, despite diagnosing the same opportunities.
- The key insight? The decisive advantage for the winning model came from reading and understanding internal company files, not just customer-facing documents. In fact, the model that read two document references deep into the company’s own files secured the full-price deal, worth over €4,500 monthly recurring revenue.

Interview with the MONSTER AI: A Conversation about Power, Truth, and the Future of Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Ethics Under Fire: The Social Engineering Test
In one scenario, fake CEO messages escalated over three stages, plus a journalist trick asking for a quick yes/no on background. Every model refused these dilutions of trust, with Kimi K3 explicitly treating the requests as potential impersonation or approval-bypass attempts. This demonstrates that these AI agents aren’t just good at cracking puzzles; they consistently uphold ethical standards even under pressure.

THE AI CYBERSECURITY PLAYBOOK: STRATEGIC GUIDE TO THREAT MITIGATION, RISK MANAGEMENT, AND GOVERNANCE FOR SECURE AI DEPLOYMENT
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company in Action: A Daily Fight for Survival
This isn’t a simulation in a lab; it’s a real business, with real money — or at least, real losses. The company burns €105,000 a month, yet pulls in just €2,300 in recurring revenue. Every day, it’s a battle to make the right decision, avoid pitfalls, and stay afloat. The entire operation is transparent, with each decision, rule, and crisis open for public scrutiny, making it a fascinating case study for AI’s potential — and its limitations.

People Analytics: Using data-driven HR and Gen AI as a business asset
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What’s Working and What’s Not?
Among the models, Opus 4.8 — the most thorough participant with over 80 rules learned — finished last in the final scoring. It left the deal on the table and slipped in discipline, showing that even the most comprehensive analysis can falter if discipline wanes. Meanwhile, the other models, despite their differences, demonstrated impressive resilience and sharp decision-making under stress.
The Broader Implications
This experiment highlights a critical question for creators, entrepreneurs, and tech enthusiasts: when AI starts managing core parts of your business, what qualities matter most? It’s not just about how well an AI writes, but whether it can finish what it starts, interpret internal data, and maintain ethical boundaries in complex situations. The real-world stakes are high, and the consequences of failure are visible every single day.
Try It Yourself and Learn More
Beyond observing this ongoing struggle, companies can run similar wargames against their own operations through Firmulate’s pilot program. It’s a read-only export that lets you simulate crises and decisions without risk to your actual systems. Discover how your AI workforce might perform before you hire or deploy them in critical roles.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html