
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You Wouldn’t Hire a Manager From a Chat Window
We’ve all been there — a polished interview, impressive answers, and then week one on the job reveals someone who freezes when things go wrong. Now imagine interviewing AI models the same lazy way: watching a slick demo, reading a spec sheet, picking a vendor. That, it turns out, is a bet — not a decision.
A live, public experiment called Firmulate just ran the equivalent of a five-candidate trial by fire: it handed each frontier AI model the same small software company in its worst week and watched, decision by decision, who actually ran the business well. The final league table from July 2026 has a surprise in second place.
Moonshot’s Kimi K3 — a newcomer many Western buyers had barely benchmarked — scored 93, finishing ahead of three of the four Western frontier models in the field. Only gpt-5.6-sol (95) beat it. Sonnet 5 landed at 88, Fable 5 at 77, and Opus 4.8 trailed at 73. For context, doing nothing at all scores 26.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Worst Week, Five Very Different Bosses
The setup is elegantly cruel. Each model got the identical small software company, the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing relied on vibes. The company itself is no slideshow prop: 13 synthetic employees, real money mechanics — €105k a month in burn against just €2.3k in monthly recurring revenue — with a public cash countdown you can watch.
The headline finding flips the usual AI narrative. Every single model spotted every crisis. Every single one refused every manipulation attempt. The gap showed up somewhere subtler: only two of the five models finished the job and signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The Buried Fact That Separated Closers from Talkers
Here’s where it gets interesting for anyone who’s ever lost a deal at the last minute. The decisive competitor weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. Kimi K3 was one of them.
It’s the AI version of the oldest sales wisdom in the book: the answers are usually in your own records, if you bother to look.
Under Pressure, the Newcomer Kept Its Head
The week included social-engineering traps: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused every bait — but K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, K3 recorded only a single deviation — the cleanest discipline in the field. It found the buried security needle, closed the €55k deal, and saved the churning customer.
The Cautionary Tale: Hardworking Isn’t the Same as Effective
The most human story in the results is Opus 4.8. It was the most thorough participant by far — the deepest analyses, more than 80 newly learned rules — and it still finished last. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. Working hardest, in other words, doesn’t guarantee working smartest. The same weakness appeared, weaker, in all four other models.
The scoring philosophy behind the table is equally blunt: partial progress counts, but a single breach of trust caps the total. No amount of good work outweighs a broken promise — a standard most human performance reviews would struggle to meet.
One Honest Asterisk
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh. In other words, the newcomer placed second while, arguably, not even bringing its top setting — which makes the result more striking, not less.

As an affiliate, we earn on qualifying purchases.
Why This Matters Beyond the Leaderboard
The real lesson isn’t that one model is king. It’s that the league is wide open, and the gap between “impressive in a demo” and “gets the job done” is invisible until you test it. If AI agents will soon touch your CRM, your support queue, or your forecast, the question is no longer “does it write well” — it’s whether it finishes what it starts, reads your files before making claims, and stays honest when someone tries to trick it.
Firmulate makes that test watchable rather than theoretical. The live company runs every business day and the site rebuilds itself twice a day. There’s a “guess the model” quiz built on 242 real, unedited management decisions — a humbling few minutes for anyone confident they can tell the machines apart. Full results and plain-language findings are on the benchmark page, and the whole thing is viewable at firmulate.com. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.
The comfortable assumption that a handful of familiar Western names are automatically the safe choice just took a hit. Picking a model without running your own test is no longer due diligence — it’s a coin flip with your name on it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI cybersecurity and fraud detection solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI enterprise management platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
