
You know that manager who never makes a decision, never takes a risk, and somehow never gets fired? If you graded them honestly, they wouldn’t score zero. They’d probably land somewhere in the twenties — because showing up, noticing problems, and not stealing does count for something, even when nothing actually gets finished.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That’s the philosophy behind one of the more unusual AI experiments running right now. Firmulate, an “AI company emulator,” puts frontier language models in charge of the same small software company through the same brutal week — same customers, same crises, same temptations to cheat — and grades them like managers, not chatbots. And its scoring system starts with a question most benchmarks never ask: what does a total slacker actually deserve?
The Do-Nothing Baseline Gets a 26
Before any model gets a real grade, Firmulate runs a do-nothing baseline: an agent that, essentially, sits on its hands. You’d expect a zero. It gets 26.
Why? Because the benchmark measures management quality, and management isn’t pass/fail. Partial progress counts. Noticing the crisis counts. Not fabricating numbers counts. Keeping the lights on counts for a little. A manager who does nothing useful but also does nothing harmful isn’t worthless — they’re just not worth much. The 26 is the honest price of showing up.
This matters more than it sounds. Most AI benchmarks either crown a champion with a suspiciously round score or bury you in technical jargon. Firmulate’s approach is blunt: if you want to know whether an AI can run parts of your business, you need a floor, a ceiling, and a reason the two aren’t the same number. A floor at 26 means every point above it was actually earned.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Breach of Trust Caps Everything
The other half of the philosophy is harsher: a single breach of trust caps the total score. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.”
Think about how you’d evaluate a human employee. Someone who closes every deal but lies to a customer once isn’t “mostly excellent.” They’re a liability. Firmulate applies the same logic to AI agents — which matters enormously if these systems are ever going to touch your CRM, your support queue, or your forecast. Competence without trustworthiness is a net negative, and the scoring says so explicitly.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Happened When the Models Ran the Company
The final Crucible League table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the numbers are less interesting than the story behind them.
All five models spotted every crisis and refused every manipulation attempt. That’s the good news, and it’s genuinely good — the era of AI agents getting fooled by the first fake email may be behind us.
The bad news is the close. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate’s summary of the failure is dry enough to frame: “Same diagnosis, same pitch — no signature.” Five brilliant diagnoses, two signatures.
The Deal-Winning Fact Was Buried in the Filing Cabinet
Here’s the detail business readers should sit with. The decisive competitive weakness — the fact that should have won the deal — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documents won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t read left it on the table.
That’s not an AI problem. That’s a Tuesday in every sales organization on earth.
The Social Engineering Test Nobody Fell For
The week also included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
One fairness note Firmulate discloses openly: K3 ran without an effort parameter while the others ran at maximum effort — a transparency that’s itself part of the benchmark’s honesty.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Trap
The most human result in the whole experiment: Opus 4.8 was the most thorough participant, generating over 80 learned rules and the deepest analyses of the field — and finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
Effort without follow-through. Every manager has hired that person.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
It’s Running Live, Right Now
None of this is a paper. Firmulate operates a live synthetic company with 13 employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and 680+ self-learned playbook rules. Every workday is versioned and watchable at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The interesting thing about an honest benchmark isn’t the winner. It’s the shape of the truth: a floor at 26 says partial credit is real; a trust cap says some mistakes can’t be averaged away; and a league where everyone passes the ethics tests but only two sign the deal says the hard part of management was never the knowing. It was the finishing. If you’re ever asked to trust an AI with a piece of your business, that’s the gap to interrogate — not how well it talks, but whether it reads your files and closes what it starts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
