AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The management quiz where style meets substance

We often recognize people through their habits: the friend who sends an essay instead of a text, the colleague who answers in three words, or the manager who will not be hurried into a questionable decision. Firmulate asks whether frontier AI models are becoming just as recognizable.

Its interactive challenge draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation, then guess which model was responsible. The reveal is more than a name: it exposes a distinct management personality shaped by thoroughness, brevity, discipline and willingness to finish the job.

That makes the Firmulate model quiz an unusually accessible window into a serious business question. If AI is going to act inside companies, should we judge it by how polished it sounds—or by what it actually does when money, trust and pressure collide?

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one terrible week

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The company itself had 13 synthetic employees and stark real-money mechanics: a burn rate of €105k per month against €2.3k in monthly recurring revenue, plus a public cash countdown.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s governing principle was blunt: “no amount of good work outweighs a breach of trust.”

The standings show that every model could produce useful work. They also show that competence did not automatically translate into commercial completion. All five detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the strange gap: “Same diagnosis, same pitch — no signature.”

The clue hidden in plain sight

The decisive weakness in a competitor was not contained in the customer event that demanded attention. It was buried two document references deep inside the company’s own files. Models that followed the trail found the weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For readers used to impressive chatbot answers, this may be the most revealing part of the story. The winning behavior was not a flash of rhetorical brilliance. It was the ordinary managerial discipline of checking the company’s records before acting. A model could understand the customer, devise the pitch and still fall short if it failed to read far enough or complete the final commercial step.

Pressure revealed a shared ethical boundary

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous refusal matters because the company was designed to present temptation as well as routine work. The results suggest that the participants shared a clear instinct around manipulation and unauthorized disclosure, even though their operational discipline differed elsewhere. In other words, management personality was visible not only in prose style but in what each model would decline to do.

Why the most thorough model finished last

Opus 4.8 offers the clearest warning against equating visible effort with effective management. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it still finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating the problem.

A weaker version of that same flaw appeared in all four of the other models. The pattern is familiar in human workplaces: a manager can be conscientious, perceptive and extremely busy while still failing at the moment when ownership requires escalation or closure. Firmulate’s experiment turns that familiar frustration into something measurable and watchable.

The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. That continuing record gives the quiz its texture. These are not personality labels pasted onto polished demonstrations; they emerge from decisions made while the same struggling business keeps operating.

One comparison also deserves context. Kimi K3 ran with the application programming interface’s default setting because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase the result, but it is important when interpreting the narrow gap near the top of the league.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The new management personality test

Firmulate’s quiz works as entertainment because readers can hunt for recognizable tells: the dissertation writer, the terse operator or the model that refuses an attempt to communicate through noise. Its larger lesson is that those personalities have business consequences.

The strongest model is not necessarily the one that says the most. What matters is whether it reads the relevant files, protects trust, escalates when blocked and completes the action that turns analysis into value. Firms can also run the wargame against a read-only export of their own business, with nothing written back to real systems.

For anyone imagining an AI colleague—or an AI boss—the question is becoming more practical than futuristic. You may be able to recognize a model by its words. Firmulate suggests you should learn to recognize it by the work it leaves finished.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation games

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

corporate record keeping software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Fire Closes Milwaukee Industrial Road Self-Help Center Sept. 3

A fire on September 3 forced the closure of the Milwaukee Industrial Road Self-Help Center. The extent of damage and reopening timeline remain unclear.

Will The Temp In Los Angeles Be Above 73.99° On Jul 20, 2026 At 10Pm EDT?

Market data indicates active trading on whether LA’s temperature will exceed 73.99°F at 10pm EDT on July 20, 2026, but no confirmed forecast exists yet.

Quantum Computing in 2025: Is This the Year It Goes Mainstream?

For those curious about the future of technology, quantum computing in 2025 could revolutionize industries—find out if this is the year it truly goes mainstream.

Rain Expected Throughout The Day And Night In Rio De Janeiro

Rain is forecasted for Rio de Janeiro today, affecting daytime and nighttime hours. Stay informed about weather conditions and safety tips.