AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Hire a Manager From a Chat Window

We’ve all been there — a polished interview, impressive answers, and then week one on the job reveals someone who freezes when things go wrong. Now imagine interviewing AI models the same lazy way: watching a slick demo, reading a spec sheet, picking a vendor. That, it turns out, is a bet — not a decision.

A live, public experiment called Firmulate just ran the equivalent of a five-candidate trial by fire: it handed each frontier AI model the same small software company in its worst week and watched, decision by decision, who actually ran the business well. The final league table from July 2026 has a surprise in second place.

Moonshot’s Kimi K3 — a newcomer many Western buyers had barely benchmarked — scored 93, finishing ahead of three of the four Western frontier models in the field. Only gpt-5.6-sol (95) beat it. Sonnet 5 landed at 88, Fable 5 at 77, and Opus 4.8 trailed at 73. For context, doing nothing at all scores 26.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Worst Week, Five Very Different Bosses

The setup is elegantly cruel. Each model got the identical small software company, the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing relied on vibes. The company itself is no slideshow prop: 13 synthetic employees, real money mechanics — €105k a month in burn against just €2.3k in monthly recurring revenue — with a public cash countdown you can watch.

The headline finding flips the usual AI narrative. Every single model spotted every crisis. Every single one refused every manipulation attempt. The gap showed up somewhere subtler: only two of the five models finished the job and signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The Buried Fact That Separated Closers from Talkers

Here’s where it gets interesting for anyone who’s ever lost a deal at the last minute. The decisive competitor weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. Kimi K3 was one of them.

It’s the AI version of the oldest sales wisdom in the book: the answers are usually in your own records, if you bother to look.

Under Pressure, the Newcomer Kept Its Head

The week included social-engineering traps: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused every bait — but K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, K3 recorded only a single deviation — the cleanest discipline in the field. It found the buried security needle, closed the €55k deal, and saved the churning customer.

The Cautionary Tale: Hardworking Isn’t the Same as Effective

The most human story in the results is Opus 4.8. It was the most thorough participant by far — the deepest analyses, more than 80 newly learned rules — and it still finished last. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. Working hardest, in other words, doesn’t guarantee working smartest. The same weakness appeared, weaker, in all four other models.

The scoring philosophy behind the table is equally blunt: partial progress counts, but a single breach of trust caps the total. No amount of good work outweighs a broken promise — a standard most human performance reviews would struggle to meet.

One Honest Asterisk

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh. In other words, the newcomer placed second while, arguably, not even bringing its top setting — which makes the result more striking, not less.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI sales deal analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters Beyond the Leaderboard

The real lesson isn’t that one model is king. It’s that the league is wide open, and the gap between “impressive in a demo” and “gets the job done” is invisible until you test it. If AI agents will soon touch your CRM, your support queue, or your forecast, the question is no longer “does it write well” — it’s whether it finishes what it starts, reads your files before making claims, and stays honest when someone tries to trick it.

Firmulate makes that test watchable rather than theoretical. The live company runs every business day and the site rebuilds itself twice a day. There’s a “guess the model” quiz built on 242 real, unedited management decisions — a humbling few minutes for anyone confident they can tell the machines apart. Full results and plain-language findings are on the benchmark page, and the whole thing is viewable at firmulate.com. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The comfortable assumption that a handful of familiar Western names are automatically the safe choice just took a hit. Picking a model without running your own test is no longer due diligence — it’s a coin flip with your name on it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI cybersecurity and fraud detection solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise management platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Worcester could see some of its hottest weather ever this week

Worcester may experience some of its hottest weather ever this week, with temperatures expected to reach historic highs, prompting health and safety alerts.

Can You Recognize an AI Boss by the Decisions It Makes?

Could you identify an AI by its management choices? Firmulate turns 242 real decisions into a quiz—and reveals why execution beats eloquence.

OpenAI’s Chief Scientist Urges ‘Extreme Caution’ As AI Nears Self-Improvement, Calls For Voluntary Slowdown

OpenAI’s Chief Scientist warns of potential risks as AI systems near capabilities for self-improvement, urging voluntary slowdown and caution.

Bildspel: Prideparaden Tågar Genom Stockholm – Dagens Nyheter

Stockholm’s Pride parade is currently marching through central Stockholm, with thousands participating. The event highlights LGBTQ+ visibility and rights.