
You wouldn’t hire a personal assistant who answered every question brilliantly but never opened your filing cabinet. You wouldn’t trust a financial advisor who gave confident advice without reading your actual bank statements. Yet that is exactly the kind of AI assistant most of us are inviting into our businesses and lives — polished talkers whose homework habits we’ve never actually tested.
A public experiment called Firmulate just made that lazy-assistant problem measurable, and the results are surprisingly human: the difference between winning and losing a €55,000 deal wasn’t intelligence, charm, or effort. It was whether the AI read the files before answering.
Same company, same worst week, four different AIs
Here’s the setup, and it’s worth pausing on how clean it is. Four frontier AI models were each handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — the only variable was which model was in charge. Every decision was versioned and auditable, so there’s no arguing after the fact about who did what.
The company itself is real enough to hurt: 13 synthetic employees, real money mechanics, burn of €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown. It’s watchable live — the whole thing runs like a business reality show, with 680+ self-learned playbook rules and every workday archived.
As an affiliate, we earn on qualifying purchases.
Everybody passed the ethics test. Almost everybody flunked the homework test.
The headline finding flips the usual AI anxiety on its head. All four models spotted every crisis. All four refused every manipulation attempt — including a three-stage fake-CEO impersonation scam and a reporter’s sly “just one yes/no, on background” trick. One model, Kimi K3, even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Five out of five attempts at social engineering: refused.
Then came the €55,000 deal. Every model diagnosed the customer’s problem correctly. Every model made the right pitch. But only two actually signed the contract their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The fact that was buried two documents deep
Why did two AIs close and two freeze? The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that only surfaced if the agent actually followed the paper trail before responding.
The models that did the reading didn’t just win the deal. They won it at full price, worth +€4,583 in monthly recurring revenue. The models that skipped the file work didn’t just miss a bonus — they lost the deal automatically. Reading your files before answering turned out to be a purchase-deciding property, not a nice-to-have.
AI business decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The final league table
- 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93 points. The newcomer (from Moonshot) also closed the deal, with the cleanest discipline in the field. One caveat: K3 ran at its API default effort setting while the others ran at maximum effort — and still nearly won.
- 3. Sonnet 5 — 88 points. Closed the deal too, with a few more process slips.
- 4. Fable 5 — 77 points. Mid-table.
- 5. Opus 4.8 — 73 points. The cautionary tale — more below.
For context, a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total completely: “no amount of good work outweighs a breach of trust.” It’s a scoring philosophy most of us would recognize from real life.
As an affiliate, we earn on qualifying purchases.
The tortoise that didn’t finish
The most instructive profile is Opus 4.8. It was the most thorough participant in the entire field — over 80 self-learned rules added, the deepest analyses of anyone. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the request properly. The same weakness appeared, more mildly, in all four models. Brilliance without follow-through is a pattern, not a fluke.
If that sounds like someone you’ve worked with — or been — that’s rather the point. Firmulate measures management quality, not chat quality.

Why this matters beyond the lab
If AI agents are going to touch your CRM, your support queue, or your forecast, the useful question is no longer “does it write well?” It’s: does it finish what it starts, does it read your files first, and does it stay honest under pressure?
You can engage with this yourself. A “guess the model” quiz at Firmulate is powered by 242 real, unedited management decisions from the experiment — a surprisingly addictive way to test your own judgment against the leaderboard. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.
The lesson for the rest of us is simpler still. The winning AIs weren’t the smartest talkers. They were the ones that did their homework before opening their mouths. That standard has served humans well for centuries — and now, apparently, it scores 95 points on AI too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html