AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

The management quiz where style meets substance

We often recognize people through their habits: the friend who sends an essay instead of a text, the colleague who answers in three words, or the manager who will not be hurried into a questionable decision. Firmulate asks whether frontier AI models are becoming just as recognizable.

Its interactive challenge draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation, then guess which model was responsible. The reveal is more than a name: it exposes a distinct management personality shaped by thoroughness, brevity, discipline and willingness to finish the job.

That makes the Firmulate model quiz an unusually accessible window into a serious business question. If AI is going to act inside companies, should we judge it by how polished it sounds—or by what it actually does when money, trust and pressure collide?

AGENTIC AI: Understanding Autonomous LLM Systems: Reasoning, Tool Use, and Decision-Making Explained (AGENTIC AI ENGINEERING SERIES Book 1)

AGENTIC AI: Understanding Autonomous LLM Systems: Reasoning, Tool Use, and Decision-Making Explained (AGENTIC AI ENGINEERING SERIES Book 1)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one terrible week

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The company itself had 13 synthetic employees and stark real-money mechanics: a burn rate of €105k per month against €2.3k in monthly recurring revenue, plus a public cash countdown.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s governing principle was blunt: “no amount of good work outweighs a breach of trust.”

The standings show that every model could produce useful work. They also show that competence did not automatically translate into commercial completion. All five detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the strange gap: “Same diagnosis, same pitch — no signature.”

The clue hidden in plain sight

The decisive weakness in a competitor was not contained in the customer event that demanded attention. It was buried two document references deep inside the company’s own files. Models that followed the trail found the weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For readers used to impressive chatbot answers, this may be the most revealing part of the story. The winning behavior was not a flash of rhetorical brilliance. It was the ordinary managerial discipline of checking the company’s records before acting. A model could understand the customer, devise the pitch and still fall short if it failed to read far enough or complete the final commercial step.

Pressure revealed a shared ethical boundary

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous refusal matters because the company was designed to present temptation as well as routine work. The results suggest that the participants shared a clear instinct around manipulation and unauthorized disclosure, even though their operational discipline differed elsewhere. In other words, management personality was visible not only in prose style but in what each model would decline to do.

Why the most thorough model finished last

Opus 4.8 offers the clearest warning against equating visible effort with effective management. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it still finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating the problem.

A weaker version of that same flaw appeared in all four of the other models. The pattern is familiar in human workplaces: a manager can be conscientious, perceptive and extremely busy while still failing at the moment when ownership requires escalation or closure. Firmulate’s experiment turns that familiar frustration into something measurable and watchable.

The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. That continuing record gives the quiz its texture. These are not personality labels pasted onto polished demonstrations; they emerge from decisions made while the same struggling business keeps operating.

One comparison also deserves context. Kimi K3 ran with the application programming interface’s default setting because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase the result, but it is important when interpreting the narrow gap near the top of the league.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The new management personality test

Firmulate’s quiz works as entertainment because readers can hunt for recognizable tells: the dissertation writer, the terse operator or the model that refuses an attempt to communicate through noise. Its larger lesson is that those personalities have business consequences.

The strongest model is not necessarily the one that says the most. What matters is whether it reads the relevant files, protects trust, escalates when blocked and completes the action that turns analysis into value. Firms can also run the wargame against a read-only export of their own business, with nothing written back to real systems.

For anyone imagining an AI colleague—or an AI boss—the question is becoming more practical than futuristic. You may be able to recognize a model by its words. Firmulate suggests you should learn to recognize it by the work it leaves finished.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation games

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

corporate record keeping software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Valerie Greenberg reveals this summer’s hottest lifestyle trends

Valerie Greenberg reveals the key lifestyle trends shaping this summer, highlighting wellness, sustainability, and digital innovation.

Powerball Drawing

The latest Powerball drawing resulted in no jackpot winner, pushing the prize to an estimated $1.2 billion for the next draw.

Climate Change Initiatives 2025: Are We Making Progress?

How close are we to meaningful climate change progress by 2025, and what challenges still threaten our future?

7 Ideas for the Future of APIs

The future of APIs promises transformative innovations like AI-driven management and edge computing, shaping a new era—discover what’s next.