
For people who rely on assistive technology, accessibility depends on more than whether a system can recognize a request. It also has to respond reliably when the stakes rise. The same question faces businesses bringing AI agents into customer support, sales and operations: can they spot trouble, follow the right process and finish the job?
Get everyday helpers delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts that question to a live, watchable experiment. Its AI-run company faces real-money mechanics and business crises, offering a look at how models behave beyond a polished chat exchange. Follow the live company.
A shared week under pressure
In the final Crucible League, published in July 2026, five models faced the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was less about spotting a crisis than acting on what the analysis showed. All models identified every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary is pointed: “Same diagnosis, same pitch — no signature.”
The clue was already in the files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For businesses, that is a reminder that an agent may need to connect information across ordinary records before it can make a consequential decision.
Firmulate also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as execution
Opus 4.8 produced the most thorough participation, adding 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. The experiment therefore raises a practical question for organizations: will an agent act within its authority, and will it escalate when it cannot?
There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The rankings are useful context for this particular experiment, not a universal verdict on every deployment.
From watching to a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. A separate quiz uses 242 real, unedited management decisions and asks visitors to guess which model made them.
For an enterprise, the next step is to test agents against its own conditions. Firmulate’s pilot uses a read-only export of a company’s business to stage crisis scenarios and produce a board report with model rankings and weaknesses in its playbooks. Nothing writes back to real systems. That makes the exercise a way to examine decisions before granting an agent operational reach.

Put the decisions under scrutiny
AI agents may recognize a problem and still fail to close the loop. Firmulate’s experiment highlights the value of testing whether models can use scattered evidence, preserve trust and escalate when they hit a boundary. Enterprises can explore a pilot against a read-only business export at firmulate.com/pilot.html. To discuss a pilot, contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
