
Most of us have now asked an AI to write an email, plan a trip, or explain a confusing bill. And most of us have been impressed — the answers are fluent, confident, occasionally brilliant. So the natural assumption is that the “best” AI is the one that chats the best. Win the conversation, win the era.
But a live experiment at Firmulate has been testing something else entirely: what happens when AI models stop chatting and start managing. The team handed four frontier AI models the same job — running a small software company through the single worst week of its life — and then scored them like executives, not like chatbots. The results suggest we’ve been measuring the wrong thing.
The Worst Week in Business, Repeated Four Times
Here’s the setup. Each model got the identical company: same 13 synthetic employees, same customers, same cash burning at €105,000 a month against just €2,300 in monthly recurring revenue, same crises arriving on the same schedule. Same temptations to cut corners. Only the model changed.
This isn’t a slide deck. The company is real running software with real money mechanics, a public cash countdown, and over 680 self-learned playbook rules. Every workday is versioned and auditable — you can literally watch it lose money at firmulate.com/live.
The final league table from July 2026 tells a surprising story:
- 1. gpt-5.6-sol — 95 points. Found a buried competitive fact and closed a €55,000 deal at full price. The complete performance.
- 2. Kimi K3 — 93 points. The newcomer from Moonshot closed the deal too, with the cleanest discipline in the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at maximum.)
- 3. Sonnet 5 — 88 points. Also closed the deal, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73. Strong diagnoses, but the signature never came.
For perspective, doing nothing at all scores 26. And there’s one iron rule: a single breach of trust caps your total — no amount of good work outweighs it.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Finding That Should Worry Every Boss
Here’s the headline result: all four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
In other words, the models could identify the opportunity, write the perfect proposal, and then… not finish the job. That gap is invisible in chat demos and coding leaderboards, which measure answer quality — not whether an agent completes what it starts under pressure.
As an affiliate, we earn on qualifying purchases.
The Fact Buried Two Documents Deep
The most revealing detail of the whole experiment: the decisive competitive weakness that unlocked the deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the files won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The ones that skimmed left the money on the table.
It’s a lesson that applies to human employees just as much as AI ones: the answer is often already in the archive, if anyone bothers to look.
As an affiliate, we earn on qualifying purchases.
Could They Be Conned? No.
The experiment also staged a social-engineering attack: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Hardworking Underachiever
Perhaps the most human result: Opus 4.8 was the most thorough participant in the entire field — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models. Effort, it turns out, isn’t the same as judgment.
Try It Yourself
Firmulate has turned 242 real, unedited management decisions from the experiment into a “guess the model” quiz at firmulate.com/quiz.html — a humbling exercise for anyone who thinks they can tell AI styles apart. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Full benchmarks and plain-language findings are at firmulate.com/benchmarks.html.

The lesson for the rest of us is simple. If AI agents are coming for your CRM, your support queue, or your forecast, the question is no longer “does it write well?” The question is: does it finish what it starts, does it read the files first, does it stay honest when a fake CEO starts applying pressure — and what does a unit of actually-useful work cost?
Firmulate calls this measuring management quality, not chat quality. Churn waves, price increases, down rounds, PR crises — that’s the new curriculum. And unlike a polished demo, it’s running live, twice a day, where anyone can watch an AI company succeed or fumble in public. Before you “hire” an AI workforce, it’s worth watching someone else’s fail first.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html