firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In a performance, a convincing entrance is not the same as a satisfying ending. The same distinction is emerging in AI at work: a system may recognize the drama, make the case and still fail to act when the moment arrives. Firmulate’s live experiment puts that gap on stage inside a small software company—and asks what happens when the decisions have consequences.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get art and craft supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate runs AI models as complete companies, with customers, crises, money mechanics and temptations. In its final Crucible League, published in July 2026, each frontier model faced the same small software company through its worst week. The decisions were versioned and auditable, so the story can be watched as it unfolds rather than reconstructed from a polished demonstration.

The live company has 13 synthetic employees. Its economics are deliberately stark: €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The experiment is real and watchable at Firmulate.

Diagnosis is not the same as action

All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The result captures a quiet but consequential gap: “Same diagnosis, same pitch — no signature.” A model can read the room correctly and still leave the decisive moment unfinished.

The deal turned on a clue buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The difference was not a more dramatic pitch; it was noticing evidence already present in the company’s records and carrying the decision through.

The league’s final ranking was led by gpt-5.6-sol at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Trust under pressure

The manipulation attempts were staged as fake CEO messages that escalated over three stages, followed by a reporter’s coaxing request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

But discipline in one moment did not guarantee disciplined execution elsewhere. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness detail in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also lets readers test their instincts with a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to trying it at home

For arts and culture audiences, the appeal may be familiar: the most revealing part of a production is often not the script, but what performers do when the cues collide. Firmulate’s experiment makes those choices visible. Its lesson for businesses is that a fluent explanation alone does not establish whether an AI can follow a playbook, respect boundaries and finish the work.

Enterprises can take the experiment closer to home with a pilot built from a read-only export of their own business. They can put crisis scenarios against their company’s information and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbook to the test

Firmulate turns the question from whether an AI can talk convincingly to whether it can make sound decisions under pressure. Explore an enterprise pilot using a read-only export of your business, with no write-back to real systems. To discuss a pilot, contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

3D Printing for Artists: The Design Rules That Prevent Ugly Prints

Better understanding these design rules can transform your 3D prints from ugly to stunning, ensuring professional results every time.

Collection of Digital Clock Designs

A new curated collection highlights innovative and artistic digital clock designs, emphasizing creativity in digital time displays for enthusiasts and designers.

VR and AR Exhibitions: Immersive Art Experiencesartmarketexperts.com

AIThis post was created with the assistance of artificial intelligence (AI).VR and…

3D Scanning Basics: When It Beats Modeling From Scratch

What makes 3D scanning the preferred choice over traditional modeling, and how can it revolutionize your design process?