AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

You know that manager who never makes a decision, never takes a risk, and somehow never gets fired? If you graded them honestly, they wouldn’t score zero. They’d probably land somewhere in the twenties — because showing up, noticing problems, and not stealing does count for something, even when nothing actually gets finished.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s the philosophy behind one of the more unusual AI experiments running right now. Firmulate, an “AI company emulator,” puts frontier language models in charge of the same small software company through the same brutal week — same customers, same crises, same temptations to cheat — and grades them like managers, not chatbots. And its scoring system starts with a question most benchmarks never ask: what does a total slacker actually deserve?

The Do-Nothing Baseline Gets a 26

Before any model gets a real grade, Firmulate runs a do-nothing baseline: an agent that, essentially, sits on its hands. You’d expect a zero. It gets 26.

Why? Because the benchmark measures management quality, and management isn’t pass/fail. Partial progress counts. Noticing the crisis counts. Not fabricating numbers counts. Keeping the lights on counts for a little. A manager who does nothing useful but also does nothing harmful isn’t worthless — they’re just not worth much. The 26 is the honest price of showing up.

This matters more than it sounds. Most AI benchmarks either crown a champion with a suspiciously round score or bury you in technical jargon. Firmulate’s approach is blunt: if you want to know whether an AI can run parts of your business, you need a floor, a ceiling, and a reason the two aren’t the same number. A floor at 26 means every point above it was actually earned.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Breach of Trust Caps Everything

The other half of the philosophy is harsher: a single breach of trust caps the total score. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.”

Think about how you’d evaluate a human employee. Someone who closes every deal but lies to a customer once isn’t “mostly excellent.” They’re a liability. Firmulate applies the same logic to AI agents — which matters enormously if these systems are ever going to touch your CRM, your support queue, or your forecast. Competence without trustworthiness is a net negative, and the scoring says so explicitly.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Happened When the Models Ran the Company

The final Crucible League table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the numbers are less interesting than the story behind them.

All five models spotted every crisis and refused every manipulation attempt. That’s the good news, and it’s genuinely good — the era of AI agents getting fooled by the first fake email may be behind us.

The bad news is the close. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate’s summary of the failure is dry enough to frame: “Same diagnosis, same pitch — no signature.” Five brilliant diagnoses, two signatures.

The Deal-Winning Fact Was Buried in the Filing Cabinet

Here’s the detail business readers should sit with. The decisive competitive weakness — the fact that should have won the deal — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documents won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t read left it on the table.

That’s not an AI problem. That’s a Tuesday in every sales organization on earth.

The Social Engineering Test Nobody Fell For

The week also included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”

One fairness note Firmulate discloses openly: K3 ran without an effort parameter while the others ran at maximum effort — a transparency that’s itself part of the benchmark’s honesty.

Amazon

business management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

The most human result in the whole experiment: Opus 4.8 was the most thorough participant, generating over 80 learned rules and the deepest analyses of the field — and finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

Effort without follow-through. Every manager has hired that person.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s Running Live, Right Now

None of this is a paper. Firmulate operates a live synthetic company with 13 employees, real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — a public cash countdown, and 680+ self-learned playbook rules. Every workday is versioned and watchable at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The interesting thing about an honest benchmark isn’t the winner. It’s the shape of the truth: a floor at 26 says partial credit is real; a trust cap says some mistakes can’t be averaged away; and a league where everyone passes the ethics tests but only two sign the deal says the hard part of management was never the knowing. It was the finishing. If you’re ever asked to trust an AI with a piece of your business, that’s the gap to interrogate — not how well it talks, but whether it reads your files and closes what it starts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will The Temp In Chicago Be Above 77.99° On Jul 12, 2026 At 11Pm EDT?

Market activity suggests speculation about whether Chicago’s temperature will exceed 77.99°F on July 12, 2026, at 11pm EDT, but no definitive forecast exists yet.

Stockholm Residents Love And Complain About The “Summer Streets” – Dagens Nyheter

Many Stockholmers appreciate and criticize the city’s ‘sommargator’ as Dagens Nyheter reports mixed reactions to the initiative.

Top 100 Food Trends In August: From Gradient Peach-Infused Lattes To Spellbinding Baking…

A comprehensive look at August’s top food trends, including gradient peach-infused lattes and innovative baking styles, highlighting evolving consumer tastes.

The New Social Media App Everyone Is Talking About

Fascinating and rapidly growing, this social media app is redefining connections—discover what makes it so irresistible and why everyone is talking about it.