
When a company’s struggle becomes the content
Most build-in-public stories showcase launches, milestones and carefully selected lessons. Firmulate exposes something less comfortable: a software company with 13 synthetic employees, real money mechanics and a widening gap between revenue and spending. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the stakes visible.
That makes the experiment compelling beyond the technology itself. It has the rhythm of a continuing business story: decisions are made, mistakes have consequences and survival is not guaranteed. Every workday is versioned, giving visitors a chance to follow the company’s behavior rather than merely read its claims. The operation can be watched live, while a separate quotes page shows what its synthetic employees actually say.

Data-Driven Decision-Making for Business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company learning in public
Firmulate has accumulated more than 680 self-learned playbook rules. That growing body of experience gives the public experiment continuity: the company is not simply producing isolated answers but carrying lessons from one workday into the next.
The commercial pressure matters because it turns abstract questions about artificial intelligence into ordinary management questions. Does the company notice a crisis? Does it read its own files before acting? Will it complete a sale after doing the analysis? Can it resist pressure to break trust? Those questions become sharper when the business is visibly spending far more than it earns.
The worst week, repeated under equal conditions
The final Crucible League results from July 2026 provide a controlled comparison. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
GPT-5.6-sol finished first with 95, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline received 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust”.
The encouraging finding was that every model spotted every crisis and refused every manipulation attempt. The unsettling finding was that only two signed the €55,000 deal their own work had earned. As Firmulate summarized the gap: “Same diagnosis, same pitch — no signature”.
The decisive information was already inside the company
The difference between analysis and commercial success came down to a buried fact. A decisive competitor weakness was located two document references deep in the company’s own files rather than in the customer event. Models that followed the references found it and won the deal at full price, adding €4,583 in monthly recurring revenue.
For a general business audience, that lesson is more significant than whether a model can write a polished message. Useful work often depends on patient reading, connecting internal information and finishing an action after the correct conclusion has been reached.
Pressure did not produce a breach
The models also faced fake messages from the chief executive that escalated across three stages, plus a reporter attempting to obtain “just one yes/no, on background”. All 5 of 5 refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s result carries a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That caveat does not erase its second-place finish, but it belongs beside the result.
Thoroughness was not enough
Opus 4.8 was the most thorough participant. It produced 80 additional learned rules and the deepest analyses, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
That contrast is central to Firmulate’s public narrative. Extensive reasoning can coexist with incomplete execution. A company fighting for survival cannot pay its bills with analysis that stops just before the decisive action.


The Lean Startup: How Today's Entrepreneurs Use Continuous Innovation to Create Radically Successful Businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why this experiment is worth following
Firmulate turns AI management into an observable business story. Its 13 synthetic employees must operate under a stark financial constraint, preserve trust and convert good judgment into completed work. The public countdown ensures that none of those questions feels theoretical.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. But the live company remains the clearest demonstration: a running experiment in whether synthetic workers can learn, resist manipulation and finish the work that keeps a business alive.
For readers accustomed to founder quotes and success stories, the attraction is unusually direct. Firmulate is publishing the difficult middle—the choices, missed opportunities and daily material produced while a company openly fights for survival.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI for Small Business: From Marketing and Sales to HR and Operations, How to Employ the Power of Artificial Intelligence for Small Business Success (AI Advantage)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.