
Interior designers know a room can look polished while hiding a practical problem: a door that won’t open fully, a walkway that pinches, or a beautiful finish that fails under daily use. Businesses face a similar test with AI. A convincing answer is not the same as a decision that holds up when customers, money and pressure are involved.
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts AI models through a live, watchable company simulation. Its enterprise pilot takes that idea to a business’s own world: test crisis scenarios against a read-only export, then review what the models did and where the company’s playbooks proved weak.
A company’s worst week, on repeat
In the final Crucible League, dated July 2026, frontier models faced the same small software company, the same customers, the same crises and the same temptations. Each decision was versioned and auditable. The final ranking put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s standard treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap, as the experiment puts it, was “Same diagnosis, same pitch — no signature.” For a business considering AI agents, recognizing the right move and completing it are different tests.
The clue was already in the company’s files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder that a business’s crucial knowledge may be scattered across its existing materials, and that a system’s performance depends on whether it brings that context to bear.
The pressure test included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of restraint matters, but the experiment also looked at whether models followed through on work they had identified as valuable.
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. The comparison is a reminder that analysis, initiative and disciplined execution do not necessarily arrive as a package.
From watching to testing your own business
The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record for every workday. The experiment is real and watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each choice.
For enterprise teams, the next step is to test models against their own business context. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The report can turn a general question—whether AI is ready for company work—into a review of how models respond to your customers, rules and pressures.
There is one comparison caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That detail belongs alongside the results when considering the ranking.

Put the judgment to a real test
A room earns trust when it works as well as it looks. An AI system deserves the same scrutiny: can it spot trouble, respect boundaries, find relevant information and carry through on a sound decision? Firmulate’s live experiment offers a view of those questions under pressure. A pilot lets an enterprise ask them against its own business, using a read-only export and scenarios, with no write-back to real systems.
To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
