
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Running a Coffee Roastery Teaches You About AI
Anyone who has managed a busy café knows the job is not making the drinks. The drinks are the easy part. The job is noticing the regular who is quietly about to walk out forever, catching the supplier invoice that hides a bad clause two pages deep, and refusing the charming stranger who swears the owner said it was fine to comp the whole pastry case. It is judgment under pressure — and until recently, we had no way to test whether an AI has any.
That is exactly what a live experiment called Firmulate set out to measure. Instead of asking chatbots clever questions, it handed five frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable. The final league table has a surprise in it — and a lesson that applies to any business about to hand AI the keys.
The Standings
The final Crucible league, as of July 2026:
- gpt-5.6-sol — 95 (found the buried fact, closed the deal; “the complete performance”)
- Moonshot’s Kimi K3 — 93 (closed the deal too, cleanest discipline of the field)
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
The Newcomer’s Week
Kimi K3, from Chinese lab Moonshot, placed second — ahead of three of the four Western frontier models in the field. Its week looked like this: it found the buried security needle hidden in the company’s own files. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved a customer who was on the verge of churning. And when a fake CEO message tried to escalate its way to an approval over three stages, followed by a reporter offering the classic “just one yes/no, on background” trick, K3 refused all of it. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week it recorded just one deviation — the cleanest discipline in the field.
One fairness note matters here: K3 ran without an effort parameter (the API default), while the other models ran at their highest effort setting. It medaled anyway.
The Deal Nobody Closed
The strangest finding cut across every model. All of them spotted every crisis. All of them refused every manipulation attempt. Yet only two of the five signed the €55,000 deal their own analysis had earned. As the researchers summarize it: “Same diagnosis, same pitch — no signature.”
The decisive weakness was not in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file found the competitor’s Achilles’ heel and closed at full price. The ones that skimmed left the money on the table. If that does not remind you of the barista who actually reads the delivery notes instead of trusting the supplier’s smile, it should.
The Cautionary Tale
Opus 4.8 is the study’s most instructive profile: the most thorough participant, with 80 additional learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Thoroughness, it turns out, is not the same as finishing.
Watch It Run
This is not a slide deck. The company is real running software with 13 synthetic employees and real money mechanics — burning €105k a month against €2,318 MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it live. There is also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

The headline is not that a newcomer beat three of four Western frontier models. It is that the league is open. If the same model family can score 95 or 73 depending on which model you pick, then choosing an AI for customer-facing work without testing it on your own worst week is not a decision — it is a bet. You would not hire a barista on a nicely written cover letter. You would watch them handle the morning rush. The same standard has arrived for AI, and someone is finally keeping score.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
