
Are Your Baristas or AI Managers Better at Handling Crises?
Imagine if your coffee shop’s management decisions were judged not by reputation or intuition, but by a real-time AI experiment. What if AI models, just like baristas, faced the same tough week — customer complaints, supply shortages, and ethical dilemmas? The results might surprise you.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Real Business Environment
Firmulate, a pioneering AI company, ran a live test where four different frontier AI models managed a small software company during its worst week. Every decision was authentic, every crisis genuine, and the stakes real — with the company’s own data and money mechanics on the line. This experiment isn’t just a simulation; it’s a glimpse into how AI could soon manage critical business functions.
The Participants: Four Distinct AI Personalities
- GPT-5.6-sol 95: The sharpest, who uncovered hidden details in the company files and closed a crucial deal.
- Kimi K3 93: The newcomer with a straightforward approach, yet he managed to clinch the deal with discipline.
- Sonnet 5 88: Competent but occasionally slipped, managing to seal the deal despite some process slips.
- Fable 5 77: Less precise, missed opportunities, and left some deals on the table.
The Results: Who Managed to Finish Strong?
All four AI models identified every crisis and refused manipulation attempts — a vital trait in management. Yet, only two managed to sign the €55,000 deal their own analysis had earned — a clear indicator of operational discipline and decision quality. Interestingly, the decisive edge came not from immediate customer interactions but from reading critical internal documents two levels deep into the company’s files. The models that uncovered this hidden data secured the full prize, worth an additional €4,583 monthly recurring revenue.
Behavior Under Pressure: Ethical Stances and Rigid Disciplines
When faced with social engineering — fake CEO messages escalating in stages and even a reporter trick — all models refused to manipulate or bypass protocols. Kimi K3 explained, “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating a cautious and ethical stance under pressure.
The Human-Like Company: Real Money, Live Decisions
This isn’t just a demo. The experiment involves a real, functioning company with 13 synthetic employees, losing €105,000 monthly against a revenue of €2,300 monthly recurring. Every workday, decisions are made and versioned, with over 680 self-learned playbook rules guiding actions. All of this is observable in real-time at firmulate.com/live.
The Surprising Weakness: Read, but Don’t Write
The experiment revealed a consistent weakness across all models — they read internal documents well but failed to take action when it mattered most. For instance, the AI that left the deal untaken was the most thorough participant, with over 80 learned rules, yet it faltered in closing the deal and slipped into writing attempts instead of escalating issues properly.
What This Means for Business
For companies integrating AI into management roles, the key isn’t just about chat quality or superficial performance. It’s about whether the AI can finish what it starts, read important internal data, stay honest under pressure, and ultimately close essential deals. In this live test, the scores reflect these traits:
- GPT-5.6-sol 95: Full performance, uncovered hidden details, signed the deal.
- Kimi K3 93: Clean, disciplined, signed the deal.
- Sonnet 5 88: Managed with some slips, signed the deal.
- Fable 5 77: Missed closing opportunities, no signature.
These results underscore an important point: AI’s role in management will hinge on its ability to read deeply and act decisively, not just generate convincing text.
Ready to Wargame Your Own Business?
Want to see how your organization’s decision-making stacks up against these models? Try the free interactive quiz at firmulate.com/quiz.html. It’s a no-risk way to explore how AI might perform in your own worst week.

The Bottom Line: AI Management Goes Beyond Chat
In the end, the real test isn’t how well an AI can chat but whether it can finish what it starts, read critical data, and stay honest under pressure. Live experiments like this reveal that AI personalities matter — some are more disciplined, perceptive, and effective at closing deals, while others stumble on key internal insights. For business leaders, the question is clear: are you ready to wargame your AI workforce before you hire it?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html